← Latest papers
💻 computer science

Crowdsourced Multilingual Speech Intelligibility Testing

This paper addresses the lack of scalable, multilingual speech intelligibility testing by proposing a crowdsourced assessment approach, detailing its experimental design, and presenting the collection and public release of multilingual speech data alongside early results.

Original authors: Laura Lechler, Kamil Wojcicki

Published 2026-08-05
📖 5 min read🧠 Deep dive

Original authors: Laura Lechler, Kamil Wojcicki

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to understand human speech, but the robot keeps getting confused by background noise or bad connections. To fix this, engineers need to test how clear the robot's voice sounds. For decades, the only way to do this was to lock people inside a soundproof room, put on high-end headphones, and listen to words in a lab. It was like trying to judge the taste of a new soup by only letting a few professional chefs sample it in a sterile kitchen. It was accurate, but slow, expensive, and impossible to scale up quickly.

However, a new wave of "generative" audio technology is changing the game. These are smart algorithms that don't just clean up noise; they actually guess and create missing parts of speech to make it sound better. The problem is, these new tools might accidentally swap one sound for another that sounds similar but means something totally different (like turning a "f" into a "p"). To catch these sneaky mistakes, we need a massive, fast, and cheap way to test speech intelligibility—the ability to understand what is being said—across many different languages. This is where the idea of "crowdsourcing" comes in: instead of a few experts in a lab, we ask thousands of regular people on the internet to listen and tell us what they hear, turning the whole world into a giant, diverse testing ground.


In this paper, researchers Laura Lechler and Kamil Wojcicki from Cisco Systems propose a new way to test speech intelligibility by taking a classic lab test and moving it to the internet. They wanted to see if they could get reliable results by asking regular people on crowdsourcing platforms to play a simple word-guessing game, rather than hiring expensive experts in a lab.

The test they chose is called the Diagnostic Rhyme Test (DRT). Think of it like a "spot the difference" game for your ears. The computer plays a single word, like "bat," and then gives you two choices on the screen: "bat" or "pat." You just have to click the one you heard. The words are carefully chosen so they only differ by one tiny sound (like the 'b' vs. the 'p'), which helps pinpoint exactly where an algorithm might be messing up. The researchers gathered these word lists in five languages: English, Spanish, French, German, and Mandarin Chinese.

To make this work online, they had to be clever. They recorded native speakers saying these words, cleaned up the audio, and put them into a survey on a website called Prolific. They paid the participants well (more than $8 an hour) and added special "trap" questions to make sure people were actually paying attention and not just clicking randomly. They also made sure participants used headphones and had normal hearing.

The team ran four different experiments to see if this "internet lab" actually worked. First, they compared their online results with a traditional lab test using Spanish speakers. They found that while the online group scored slightly lower overall (which makes sense, since they weren't in a perfect quiet room), the online test was very good at spotting the difference between good audio and bad audio. For example, when they tested a compressed audio format (NB PCMU), both the lab experts and the online crowd agreed that it was worse than the high-quality version, even if the online crowd gave it a lower score overall.

Next, they checked if the test was consistent. They ran the same test twice with the same people and once with a completely new group. The results were almost identical, showing that the test is reliable and repeatable, just like a good science experiment should be. They also tested two different audio compression formats (AMR-WB and AMR-NB) in English. The online results matched the patterns found in old lab studies perfectly, showing that the "crowd" could tell the difference between the two codecs just as well as the experts.

Finally, they tested the system across all five languages. They found that the online test successfully detected when the audio quality dropped due to compression in English, German, Spanish, and French. Interestingly, for the Chinese test, which relies heavily on musical tones rather than just consonant sounds, the compression didn't seem to hurt intelligibility much. This makes sense because the "pitch" of the voice (which carries the tone) was preserved even when the audio quality was lowered.

The authors conclude that this crowdsourced approach is a valid, cost-effective, and fast way to test speech intelligibility. While the absolute scores might be a bit lower than in a perfect lab (perhaps because people are listening in their noisy kitchens rather than soundproof rooms), the relative results are spot on. They showed that you can use the internet to get a diverse, global pool of listeners to test new audio algorithms quickly. They have even made their word lists and audio recordings available to the public so other researchers can use them to benchmark their own work. It's a step toward making sure that as our AI voices get smarter, they also stay clear and understandable for everyone, everywhere.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →