Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions
This paper introduces a linguistically grounded, dimension-level meta-evaluation benchmark for Text-to-Speech systems to reveal that current automated evaluators, including MOS predictors and Audio-LLM judges, fail to reliably capture the full breadth of perceptual speech dimensions beyond basic acoustic quality.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to speak. You want it to sound like a real human, not a glitchy video game character. But how do you know if the robot is doing a good job? For a long time, scientists have used a "Mean Opinion Score" (MOS). Think of this like a school report card where a human listener gives the robot a single grade, like a "B" or a "4 out of 5," based on how natural it sounds. Recently, we've started using super-smart AI judges (called Audio-LLMs) to give these grades automatically, hoping they can save us the time of asking humans to listen to thousands of recordings.
The big question is: do these automatic judges actually understand why a voice sounds bad? Is it because the robot is mispronouncing a word? Is it because it's speaking too fast? Or is it because the robot sounds like a robot? If the AI judge just gives a "B" without knowing which of those specific things went wrong, it's like a teacher giving a student a bad grade on a math test but not telling them if they messed up the addition, the subtraction, or the multiplication. This paper dives into whether our current automatic judges are actually smart enough to spot these specific mistakes, or if they are just guessing based on how "noisy" the audio sounds.
The researchers at ServiceNow decided to put these automatic judges to a very specific test. They built a giant, custom-made dataset of 860 speech samples, but they didn't just record random sentences. They were like mad scientists, deliberately injecting specific types of errors into the speech. They created 10 different "flavors" of mistakes, ranging from tiny word-level errors (like saying "cat" instead of "bat") to bigger issues like speaking with the wrong emotion or having a voice that sounds physically impossible for a human to make. They then had trained human linguists listen to every single clip and label exactly what went wrong.
With this perfect "answer key" in hand, they tested four different types of automatic MOS predictors and four different Audio-LLM judges. They asked these AI models to rate the speech under different conditions: sometimes just asking for a general "naturalness" score, and other times giving them a detailed checklist of the 10 specific error types to look for.
The results were a bit of a wake-up call. The traditional MOS predictors (the ones that just give a single number) turned out to be like a smoke detector that only goes off when there's a fire, but ignores a broken window. They were great at spotting obvious, low-level glitches like robotic buzzing or weird static (what the paper calls "Human Plausibility" errors), but they were completely blind to the more linguistic mistakes. If a robot mispronounced a word or put the stress on the wrong syllable, the MOS predictor often didn't care at all, giving a high score even when the human listeners were cringing.
The Audio-LLM judges were a bit more complex. They showed some ability to spot errors, but only if you asked them the right way. When the researchers just asked them, "Is this natural?", the AI judges were inconsistent and missed a lot of specific errors. However, when the researchers gave the AI a detailed "reference guide" (a schema) listing the 10 specific dimensions to check, the judges got much better at spotting word-level mistakes. But there was a catch: if you asked the AI to check all 10 dimensions at once, it often got confused and collapsed, giving the same score to everything. It worked best when you asked it to focus on just one specific type of error at a time.
Ultimately, the paper suggests that we cannot rely on a single "naturalness" score to tell us if a Text-to-Speech system is good. The current automatic judges are like students who are great at spotting a messy desk but terrible at spotting a spelling error. They tend to focus on the sound quality of the recording rather than the actual language being spoken. The authors conclude that to truly improve these systems, we need to stop treating "naturalness" as one big mystery and start evaluating speech by looking at each specific dimension—pronunciation, stress, emotion, and speed—individually. Until we do that, our automatic judges might be giving us a passing grade for a robot that sounds like a human, but is actually saying the wrong words.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.