IASNLP 2026 gives Speech and multi-modal LLMs an easy conversational flair

What began, twenty years ago, as India’s first summer school to train future AI and language technologists, continues to excel as India’s top flagship introductory course in LLMs. IIIT Hyderabad’s LTRC-incubated IASNLP 2026 presented a fresh direction for the next generation of computational linguists and AI researchers.  

Dhwani, help me build a Dungeons and Dragons campaign! 

You might also want your LLM to explain quantum physics to you like you are a ten-year old. What if you could have that free-flowing conversation with an LLM like ChatGPT in your Indian language, instead of hunting and pecking on your keyboard to give instructions. Now picture a critical patient-doctor discussion in Telugu, with an AI model capturing its rich prosodic features like tone, emotion, rhythm, tempo, and when to interrupt, while responding like a human conversation partner.  

After ChatGPT and Gemini opened the flood gates to text-based LLMs, the next big wave is already here and it is AI-powered Spoken conversational agents. Text based AI applications are only available for literate consumers and many Indian languages do not have a script. Traditionally voice AI worked in three steps; speech to text (STT), the language model and finally text to speech (TTS). Speech LLMs learn how voices and emotions sound and respond in a voice that sounds natural. 

That is the ball park where IIIT Hyderabad’s Language Technologies Research Center (LTRC) works in Speech and multi-modal LLMs. Considered one of South Asia’s largest research centers for natural language processing and speech technologies, IIIT Hyderabad’s Language Technologies Research Center (LTRC) is also recognized as the fountainhead of large language models in India.  

IASNLP and Technology powered AI Voice Agents 
The Centre hosted its flagship event – the IIIT Advanced Summer School on Natural Language Processing (IASNLP 2026) from June 12 to 27 this year.  

As India’s very first summer school to train students, IASNLP began two decades back when Prof. Rajeev Sangal and Prof. Dipti Misra envisioned the Advanced Summer School on Natural Language Processing. Topics like AI and language technologies were largely unexplored fields.  “Many of those working on language technologies in India would have attended this legendary summer school, myself included”, smiles LTRC’s Prof. Parameswari Krishnamurthy who coordinated IASNLP 2026 with Prof. Chiranjeevi Yarra

As machine translation was the popular flavour of the era in LTRC, earlier editions were mostly application-oriented themes. This year’s choice of Speech LLMs as a multimodal theme was purely driven by market popularity. “Our intention is to build Speech LLMs for the rich diversity of Indian languages and for application in domains that impact people’s lives, like healthcare, education, finance etc”, she explains. 

Experiential Lab for handpicked candidates 
The whole multimodal concept is cleaving new pathways for research.  The 500 plus applicant pool from around the world this year featured a healthy mix of final-year undergraduate students, master’s and doctoral scholars, post-doctoral researchers, faculty members, and industry professionals from computer science, electronics and communication engineering, linguistics, cognitive sciences, as well as related disciplines in artificial intelligence, speech processing and NLP. 

Sixty-one shortlisted participants of the immersive summer school included thirty-four students, ten faculty and seventeen industry professionals, drawn from corporates like Infosys and a Mauritius-based industrialist. The Interdisciplinary team comprised of scholars from computer science, linguistics and signal processing backgrounds.   

The precursor to the main workshop was a special pre-school session on June 12 and 13, hosted by senior IIIT-H student volunteers, who presented a foundational understanding of topics, aligned with the Speech LLMs theme. The two-day event was tailor-made for those with no prior exposure to speech technologies, NLP, linguistics, or programming. These sessions before the advanced coursework began, ensured that participants from diverse academic and professional backgrounds could constructively engage with the main course curriculum. 

Through a combination of rigorous academics, practical training, collaborative projects, and interaction with leading experts, the fortnight long summer school equipped participants with the knowledge and skills required to contribute meaningfully to the evolving landscape of speech and language AI. The programme blended expert lectures, hands-on laboratory sessions, guided research projects, and mentorship from leading academics and industry professionals. Participants got to work on real-world problems in speech and language processing, giving them valuable experience in developing speech-enabled AI systems.  

Prof. Bayya Yegnanarayana

Keynote speakers Prof. Bayya Yegnanarayana, Prof. Haizhou Li from the National University of Singapore and Prof. S Umesh from IIT Madras drove a sense of urgency for cross-disciplinary research.  International and Indian domain experts from University of Essex, National Research Council-Canada, IISc, IIIT and IIT Hyderabad alongside industry professionals from leading multinational corporations like Microsoft, Qualcomm, Amazon, Oracle and Ericsson shared experiences and latest trends in their area of expertise.  

Curriculum to build a critical mass of future-ready technologists 
Speech LLMs represent a significant shift in how machines process human communication. Traditional speech technologies typically rely on separate components for Automatic Speech Recognition (ASR), NLP and Text-to-Speech (TTS). Speech LLMs unify these capabilities into a single end-to-end framework, capable of listening, understanding, reasoning, and responding directly through speech. By combining advances in large language models, self-supervised speech learning, and multimodal AI, these systems are paving the way for more natural conversational interfaces, speech-to-speech translation, intelligent voice assistants, and accessibility technologies. 

The workshop began with core concepts of speech processing and large language models before progressing to advanced topics in Speech LLMs and emerging research directions. Participants explored principles behind speech recognition, language understanding, language generation, speech synthesis, and multimodal AI systems. Equal emphasis was placed on practical implementation, allowing attendees to experiment with modern toolkits, datasets, evaluation techniques, and deployment pipelines, widely used in research and industry. 

Sessions explored state-of-the-art ASR systems, advanced Text-to-Speech synthesis, and multilingual speech technologies that support diverse languages and dialects. Participants also gained a granular understanding of LLMs including transformer architectures, pre-training methods, fine-tuning techniques, reasoning capabilities, and conversational AI systems built directly over audio. They observed how models combine audio and textual information to build conversational systems, that respond naturally to spoken input.  

The curriculum also introduced multimodal reasoning, where AI systems integrate speech with other forms of information to improve understanding and interaction. They also assessed practical applications that are increasingly transforming industries like education, healthcare, customer support, and accessibility solutions. 

An important feature of the programme was its emphasis on research. Participants undertook guided projects, with opportunities to continue their work beyond the programme under the guidance of mentors. With this project-based approach, team members got to apply newly acquired concepts while contributing to ongoing research challenges in speech and language technologies. Projects were finally analysed and evaluated by an expert jury and feedback was shared. Many will take back their learnings as a research paper while several have chosen to return as interns.  

Eye On the horizon 
As Speech LLMs rapidly redefine the future of human-computer interaction, IASNLP 2026 offers the perfect test-bed for different stakeholders in academia and industry to grapple with real-life challenges. In the final analysis, the endgame is to shape the next generation of intelligent speech systems, leveraging the strengths of IIIT Hyderabad’s research-focussed Centres in India’s first fully connected ecosystem. 

Leave a Reply

Your email address will not be published. Required fields are marked *

Next post