← Latest papers
💬 NLP

Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages

This paper introduces a controlled, multidimensional pairwise evaluation framework for multilingual Text-to-Speech (TTS) systems in 10 Indian languages, utilizing over 120,000 human comparisons to construct a reliable leaderboard, analyze perceptual trade-offs, and interpret model preferences through Bradley-Terry modeling and SHAP analysis.

Original authors: Srija Anand, Ashwin Sankar, Ishvinder Sethi, Aaditya Pareek, Kartik Rajput, Gaurav Yadav, Nikhil Narasimhan, Adish Pandya, Deepon Halder, Mohammed Safi Ur Rahman Khan, Praveen S V, Shobhit Banga, Mite
Published 2026-04-24
📖 4 min read☕ Coffee break read

Original authors: Srija Anand, Ashwin Sankar, Ishvinder Sethi, Aaditya Pareek, Kartik Rajput, Gaurav Yadav, Nikhil Narasimhan, Adish Pandya, Deepon Halder, Mohammed Safi Ur Rahman Khan, Praveen S V, Shobhit Banga, Mitesh M Khapra

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine India as a giant, bustling marketplace where millions of people prefer to shout their orders, ask questions, and tell stories out loud rather than typing them on a keyboard. This is what the authors call a "Voice-First Nation."

But here's the problem: India is like a massive pot of spices. It has hundreds of languages, and people often mix them together in the same sentence (like saying "I need a chai and a bus ticket"). Making a computer speak all these languages clearly, naturally, and without sounding like a robot is incredibly hard.

This paper is essentially a massive "Taste Test" for computer voices.

Here is the breakdown of what they did, using simple analogies:

1. The Setup: The "Voice Olympics" 🏅

Instead of just asking people, "On a scale of 1 to 10, how good is this voice?" (which is like asking a judge to rate a gymnast's score out of 10), the researchers set up a Head-to-Head Tournament.

  • The Contestants: They picked 7 of the best "Text-to-Speech" (TTS) systems currently available. Some are famous commercial giants (like Google's Gemini or ElevenLabs), and some are open-source projects built specifically for Indian languages.
  • The Judges: They recruited over 1,900 native speakers from across India. These weren't just random people; they were trained to listen critically.
  • The Arena: They created a "gym" with 5,357 different sentences. These sentences were designed to be tricky:
    • The "Tongue Twisters": Hard words to pronounce.
    • The "Code-Mix": Sentences mixing English and Indian languages (e.g., "Please book a ticket for Mumbai").
    • The "Stress Test": Sentences with numbers, math formulas, and acronyms.

2. The Game: The "Blind Taste Test" 🎧

The judges listened to two audio clips at the same time (Model A vs. Model B) without knowing which company made them.

They didn't just pick a winner. They acted like food critics rating a dish on six different ingredients:

  1. Intelligibility: Can I understand the words?
  2. Expressiveness: Does it sound happy, sad, or bored?
  3. Voice Quality: Does it sound like a human or a robot?
  4. Liveliness: Is it energetic or monotonous?
  5. Hallucinations: Did the AI make up words that weren't in the text?
  6. Noise: Is there static or hissing in the background?

3. The Results: Who Won? 🏆

After collecting over 120,000 comparisons, they used a mathematical formula (like a sports ranking system) to create a leaderboard.

  • The Champion: Gemini 2.5 Pro TTS took the gold medal. It was the most consistent winner across almost all languages and tricky sentence types.
  • The Runners-Up: ElevenLabs V3 and Sonic 3 were very close behind, often swapping places depending on the specific language.
  • The Underdog: The open-source model "Indic F5" came in last. While it tried to cover all 10 languages, it struggled to match the polish of the big commercial systems.

The Big Surprise: The researchers found that while "noise" and "hallucinations" (technical errors) are important, they aren't the main reason people pick a winner. Once the voice is clear and doesn't make mistakes, people choose the voice that sounds the most "alive" and "expressive." It's like choosing a singer: if they hit the right notes, you pick the one with the most soul.

4. The "How-To" Guide for Future Tests 📝

The paper also answers a practical question: "How many judges and how many sentences do we need to get a fair result?"

  • The Judges: You don't need thousands of people to get a stable ranking. About 200 good judges are enough to see who is clearly better than whom.
  • The Sentences: You need a diverse mix. About 1,000 different sentences covering different languages and topics are needed to make sure the ranking is fair.

The Takeaway

This paper is like a consumer report for AI voices. It tells us that:

  1. We have a way to fairly test AI voices in complex, mixed-language environments.
  2. Currently, the big commercial AI models are winning the race for Indian languages.
  3. For the future, developers shouldn't just focus on making voices "error-free"; they need to focus on making them sound human, emotional, and lively, because that's what people actually care about.

The authors are releasing all their data and test sentences to the public, hoping to help other researchers build better, more human-like voices for the world's voice-first nations.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →