← Latest papers
💬 NLP

PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech

This paper introduces PSP (Phoneme Substitution Profile), an interpretable, per-dimension benchmark that evaluates Indic Text-to-Speech systems on specific accent features like retroflexion and aspiration, revealing that current commercial and open-source leaders often fail to capture these phonological nuances despite high standard intelligibility scores.

Original authors: Venkata Pushpak Teja Menta

Published 2026-04-29
📖 5 min read🧠 Deep dive

Original authors: Venkata Pushpak Teja Menta

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a robot that can read text out loud. If you ask it to read a sentence in Hindi, Telugu, or Tamil, it might get every single word "correct" according to a spell-checker. It won't misspell a single word. But if a native speaker from that region listens to it, they might say, "That sounds like a foreigner trying to speak my language." The robot is getting the spelling right, but the accent is wrong.

This paper introduces a new tool called PSP (Phoneme Substitution Profile) to measure exactly how that accent is wrong, rather than just saying "it sounds bad" or "it sounds good."

Here is a breakdown of the paper's ideas using simple analogies:

1. The Problem: The "Perfect Spelling, Wrong Accent" Trap

Current ways of testing speech robots (Text-to-Speech or TTS) are like a spelling bee. They check:

  • Intelligibility: Did it say the right words? (Word Error Rate).
  • Naturalness: Does it sound human-ish? (MOS scores).

The Flaw: A robot can get 100% on the spelling bee but still sound like a tourist. In Indian languages, there are specific "flavor" sounds that are easy for non-natives to mess up:

  • Retroflex sounds: Tongue curls back (like a 't' or 'd' made at the roof of the mouth).
  • Aspiration: A puff of air (like a soft 'h' after a consonant).
  • Vowel length: Holding a sound longer (like "cat" vs. "caat").
  • The "Zha" sound: A unique Tamil sound that doesn't exist in English.

If a robot turns a "curled tongue" sound into a regular "flat tongue" sound, it's not a spelling error; it's an accent error. Current tools miss this.

2. The Solution: The "Six-Point Accent Report Card"

The authors created PSP, which is like a six-dimensional health check for a robot's accent. Instead of one single score, it gives you a report card with six specific grades:

  1. Retroflex Collapse Rate (RR): How often did the robot fail to curl its tongue? (Like a student who keeps forgetting to do the homework).
  2. Aspiration Fidelity (AF): Did it puff out the air when it was supposed to?
  3. Vowel Length Fidelity (LF): Did it hold the long vowels long enough?
  4. Tamil "Zha" Fidelity (ZF): Did it get that one tricky Tamil sound right?
  5. Audio Distance (FAD): How far away does the robot's voice sound from a native speaker's voice in a "sound space"? (Think of this as a distance meter).
  6. Prosodic Signature (PSD): Does the robot have the right rhythm, pitch, and melody? (Is it singing the song flatly, or with the right emotion?).

The Magic: The first four are measured by comparing the robot's sounds to a "native speaker average" (a centroid) using AI listening tools. The last two measure the overall "vibe" of the audio.

3. The Experiments: Testing the Robots

The authors tested four famous commercial robots (like ElevenLabs and Cartesia) and some open-source ones on Hindi, Telugu, and Tamil.

The Big Discoveries:

  • Difficulty Ladder: The robots found Hindi easiest, Telugu harder, and Tamil the hardest.
    • Hindi: Robots got almost perfect scores on the tricky sounds.
    • Telugu: Robots started messing up about 40% of the "curled tongue" sounds.
    • Tamil: Robots messed up about 68% of those sounds.
  • The "Spelling" vs. "Accent" Disconnect: The robot that got the best spelling scores (lowest Word Error Rate) was not always the one with the best accent.
    • Analogy: Imagine a student who gets an 'A' on a math test but fails the handwriting portion. The old grading system only looked at the math grade. PSP looks at both.
  • No Perfect Robot: No single robot was the best at everything. One robot might have great rhythm (PSD) but bad "curled tongue" sounds. Another might have great sounds but a flat, boring rhythm. You have to pick the right tool for the specific job.

4. The "Voice Prompt" Trick

The authors also tested their own robot (Praxy Voice). They found that if they gave the robot a 9-second audio clip of a real human speaking Telugu as a "reference," the robot suddenly sounded much more native.

  • Analogy: It's like a student listening to a native speaker for a few seconds before taking a test. The robot "caught the vibe" and improved its accent, even though it didn't relearn the whole language.

5. Why This Matters

The paper argues that we need to stop treating "accent" as a vague feeling. By breaking it down into these six specific dimensions, developers can see exactly where their robot is failing.

  • If the "Retroflex" score is low, they know to fix the tongue-curling sounds.
  • If the "Prosody" score is low, they know to fix the rhythm and emotion.

In short: This paper gives engineers a detailed map to navigate the tricky terrain of Indian accents, showing that getting the words right is only half the battle; getting the flavor right is the other half.

(Note: The paper explicitly states that these results are based on pilot tests and that formal correlation with human listening scores will be finalized in a future version, v2.)

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →