← Latest papers
💬 NLP

Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages

This paper proposes an orthographically-informed evaluation framework (OIWER) leveraging large language models to address the challenges of spelling variations in Indian languages, demonstrating that it provides a more accurate and human-aligned assessment of speech recognition performance compared to traditional metrics like Word Error Rate.

Original authors: Kaushal Santosh Bhogale, Tahir Javed, Greeshma Susan John, Dhruv Rathi, Akshayasree Padmanaban, Niharika Parasa, Mitesh M. Khapra

Published 2026-03-03
📖 4 min read☕ Coffee break read

Original authors: Kaushal Santosh Bhogale, Tahir Javed, Greeshma Susan John, Dhruv Rathi, Akshayasree Padmanaban, Niharika Parasa, Mitesh M. Khapra

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a teacher grading a student's spelling test. The student writes a sentence perfectly, but you mark it wrong because they used a slightly different way to spell a word that is actually acceptable in their dialect. You might say, "You missed the 'u' in 'colour'!" when the student wrote "color," or perhaps they split a compound word differently.

In the world of Automatic Speech Recognition (ASR) for Indian languages, this is exactly what's happening. Computers are trying to "listen" to human speech and turn it into text. But because Indian languages are incredibly rich, flexible, and diverse, the computers are getting unfairly harsh grades.

Here is a simple breakdown of the paper's story, using some everyday analogies:

1. The Problem: The "Rigid Grader"

For years, researchers have used a metric called WER (Word Error Rate) to grade how good these speech-to-text systems are. Think of WER as a strict, by-the-book grader.

  • The Issue: Indian languages are like a giant, flexible puzzle. You can often spell a word in five different ways, split a long word into two smaller ones, or merge two words into one, and all of them are correct.
  • The Result: The computer says, "I heard 'passbook'." The human grader wrote "pass book." The strict grader marks it as an error.
  • The Consequence: The systems look terrible on paper (high error rates), even though a human listening to the audio would say, "Hey, that sounds perfect!" It's like getting an 'F' on a test because you used a synonym instead of the exact word the teacher had in mind.

2. The Solution: The "Flexible Dictionary"

The authors (a team from IIT Madras and Sarvam AI) realized they needed a new way to grade. They created a framework called OIWER (Orthographically-Informed Word Error Rate).

Think of this as giving the grader a super-powered, flexible dictionary before they start grading.

  • Instead of just checking if the student wrote "Passbook," the dictionary tells the grader: "Wait! 'Pass book', 'pass-book', and 'passbook' are all valid. If the student wrote any of these, give them full credit."
  • They also account for tricky things like:
    • Loan words: Words borrowed from English that get spelled differently in Hindi or Tamil.
    • Sandhi: The magical way Indian languages glue words together (like how "at" + "home" becomes "athome" in speech).
    • Diacritics: The little dots and lines above or below letters that change the sound.

3. The Magic Tool: The "AI Assistant"

You might ask, "How do you make a dictionary that knows every possible way to spell a word in 22 different languages? That sounds like a million hours of work!"

The authors used Large Language Models (LLMs)—the same kind of AI that powers chatbots—to do the heavy lifting.

  • The Analogy: Imagine you need to list every possible nickname for a friend named "Robert." You could spend years asking people, or you could ask a super-smart AI: "Hey, list every way 'Robert' can be written or spoken." The AI instantly gives you "Rob, Bob, Bobby, Robbie," and even regional variations.
  • The team used this AI to generate lists of valid spellings for every word in their test data. Then, human experts just did a quick "spot check" to make sure the AI didn't get crazy.

4. The Results: A Fairer Game

When they re-graded the speech systems using this new "Flexible Dictionary" (OIWER), the results were eye-opening:

  • The "Pessimism" Vanished: The error rates dropped significantly (by an average of 6.3 points). The systems weren't actually worse; they were just being graded too harshly.
  • The Gap Closed: Previously, one model (Gemini) looked much worse than another (Canary). But when graded fairly, the gap between them shrank. It turns out they were actually much closer in performance than we thought.
  • Human Feel: The new scores matched what actual humans felt when listening to the recordings. The "computer score" finally matched the "human experience."

5. The Takeaway

This paper is essentially saying: "Stop grading Indian languages like they are rigid, standardized English."

By acknowledging that Indian languages are fluid and flexible, and by using AI to help us understand those rules, we can finally see the true potential of these speech systems. It's like finally realizing that the student didn't fail the spelling test; the test itself was just too rigid for the language being spoken.

In short: They built a smarter way to grade speech-to-text that respects the beautiful chaos of Indian languages, making the technology look as good as it actually feels to our ears.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →