← Latest papers
💬 NLP

Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages

This paper introduces Indic DiarBench, an open-access benchmark dataset featuring 108 hours of human-annotated, multi-speaker audio across all 22 scheduled Indian languages to evaluate and advance joint automatic speech recognition and speaker diarization systems.

Original authors: Deovrat Mehendale, Aditya Mehndiratta, Dhruv Rathi, Kaushal Bhogale, Mitesh M. Khapra

Published 2026-07-28
📖 4 min read☕ Coffee break read

Original authors: Deovrat Mehendale, Aditya Mehndiratta, Dhruv Rathi, Kaushal Bhogale, Mitesh M. Khapra

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking into a bustling, chaotic Indian marketplace. Hundreds of people are shouting, laughing, and talking over one another. Some are selling spices, others are haggling over prices, and a group nearby is debating politics. To a human ear, this is a rich tapestry of life, but to a computer, it's a nightmare. For years, computers have been getting very good at listening to one person speaking clearly, like a news anchor or a teacher. But real life is rarely that quiet. It's a messy, overlapping conversation where voices blend, switch languages mid-sentence, and echo off walls.

This is the world of "speaker diarization" and "speech recognition." Think of speaker diarization as a computer trying to figure out "who spoke when." It's like a referee in a noisy game trying to tag every player with a name tag as they run past. Speech recognition is the computer's attempt to write down exactly what those players said. The tricky part is that these two jobs are deeply connected. If the referee tags the wrong person, the scribe writes down the wrong words. Until now, most computer tests for this were like practicing in a silent library; they didn't prepare the computers for the loud, mixed-up, multi-language chaos of the real world, especially in India, where 22 different official languages and countless dialects swirl together.

This paper introduces a new, massive test called Indic DiarBench. The researchers didn't just build a test; they built a "chaos simulator" specifically for Indian languages. They gathered about 108 hours of real, messy audio. This isn't just one type of noise; it includes close-up recordings of people in virtual meetings, distant recordings where voices echo in large rooms, and even clips from YouTube videos where people are talking in the wild. They made sure to include all 22 scheduled languages of India, capturing everything from Hindi and Tamil to Santali and Kashmiri.

The team then put this data through a rigorous human review process. They didn't just let a computer guess; they had human experts listen, fix the transcripts, and carefully label exactly who was speaking and when, even when two people talked at the exact same time. They also handled the unique Indian habit of "code-mixing," where speakers switch between their native language and English in the same sentence.

Once the test was ready, they threw the best computer systems at it to see how they fared. They tested big commercial tools (like AWS and Azure), specialized Indian AI systems, and even the newest "multimodal" AI models that can see and hear. The results were a mix of good news and a reality check.

The study found that while some systems are getting better, the job is still incredibly hard. A specialized Indian AI system called Sarvam performed the best overall, but even it made mistakes about 16% of the time just in identifying who was speaking, and about 39% of the time in getting the words right. The big, general-purpose AI models (like GPT-4o and Gemini) showed a funny weakness: some were great at writing down the words but terrible at knowing who said them, while others were okay at identifying speakers but struggled to write the words correctly, especially in languages with fewer speakers.

The paper explicitly shows that the more people talking over each other (overlap), the worse the computers get. It also highlights that systems trained on "clean" audio fail miserably when faced with the distant, echoey, or noisy recordings found in real life. The researchers conclude that we cannot just treat "who spoke" and "what they said" as separate problems anymore; they must be solved together. While we have a new, open benchmark to help researchers improve, the paper makes it clear that we are not there yet. The computers still struggle with the vibrant, overlapping, multi-language reality of Indian conversations, and there is still a long way to go before they can listen to a crowded Indian marketplace as well as a human can.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →