← Latest papers
⚡ electrical engineering

Assessing speech quality metrics for evaluation of neural audio codecs under clean speech conditions

This paper evaluates 45 objective speech-quality metrics against subjective listening scores for 17 neural audio codecs under clean conditions, finding that neural-based metrics like ScoreQ and UTMOS achieve the highest correlation while noting that non-intrusive metrics tend to saturate at high quality levels.

Original authors: Wolfgang Mack, Nezih Topaloglu, Laura Lechler, Ivana Balić, Alexandra Craciun, Mansur Yesilbursa, Kamil Wojcicki

Published 2026-08-05
📖 4 min read☕ Coffee break read

Original authors: Wolfgang Mack, Nezih Topaloglu, Laura Lechler, Ivana Balić, Alexandra Craciun, Mansur Yesilbursa, Kamil Wojcicki

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to judge how good a song sounds after it has been squished into a tiny file to send over the internet. In the world of digital audio, this "squishing" is done by something called a codec. Think of a codec as a super-efficient chef who chops a giant, juicy steak (the original sound) into tiny, bite-sized pieces so it fits in your pocket, then tries to reassemble it later. Sometimes the reassembled steak tastes amazing; other times, it's a bit dry or weirdly textured.

For decades, engineers have used two main ways to check if the steak tastes good. The first is the Subjective Test: they invite a bunch of people to listen and rate the food on a scale of 1 to 5. It's the "gold standard," but it's slow, expensive, and requires a lot of human patience. The second way is the Objective Metric: a computer program that looks at the sound waves and tries to guess the rating without a human ever hearing it. It's fast and cheap, but for a long time, these computer judges were trained on old-school audio problems. Now, a new generation of "neural" codecs has arrived—these use artificial intelligence to recreate sound, almost like a digital artist painting a picture from scratch. The big question is: do the old computer judges still know how to grade this new AI art, or are they looking at the wrong things?

This paper sets out to answer that question by organizing a massive taste-test showdown. The researchers gathered 17 different neural audio codecs, each running at different speeds and bitrates (some as low as 1.0 kbps, others up to 8.0 kbps), and fed them 100 clean speech samples. They then asked a crowd of humans to rate the quality using a method called MUSHRA-1S, which is like a super-precise ruler capable of spotting tiny differences between high-quality sounds. Next, they ran those same sounds through 45 different computer scoring programs—some that need the original sound to compare against (intrusive) and some that just look at the result (non-intrusive).

The results were a bit like a surprise party where the expected winners didn't show up. The study found that the old-school computer judges, like PESQ and WARPQ, were okay but not great; they managed a correlation of about 0.73 with human opinions. However, the new AI-powered metrics, specifically ScoreQ, UTMOS, and Audiobox, were the clear champions, with ScoreQ hitting a correlation of 0.87. This suggests that to judge AI audio, you need an AI judge.

But there was a twist in the story. When the researchers looked closely at the very best sounds—the ones humans rated as nearly perfect—they noticed a strange glitch. The non-intrusive AI metrics (the ones that don't need the original sound) seemed to hit a "ceiling." Once the sound got really good, these metrics stopped changing their scores, even if the humans could still hear subtle differences. It's like a thermometer that stops moving once it hits 100 degrees, even if the room gets hotter. The authors suggest this happens because these metrics were trained on human ratings that were often just a simple "5 out of 5," making it hard for the computer to tell the difference between a "5" and a "5.5."

In contrast, the intrusive metrics (the ones that compare the result to the original) kept working perfectly, even at the highest quality levels, showing smaller margins of error and staying sensitive to tiny changes. The paper concludes with a practical rule of thumb: if you are testing average or slightly noisy audio, the fast, non-intrusive AI metrics like UTMOS and ScoreQ are your best friends. But if you are polishing high-fidelity audio and need to spot the tiniest imperfections, you should stick with the intrusive methods like ScoreQ Ref, which act like a magnifying glass that doesn't get blurry at the top end. The study didn't just find the best tools; it also showed us exactly where each tool stops working, helping engineers choose the right judge for the right job.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →