ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions
The paper introduces ESCUCHA, the first Spanish speech understanding benchmark featuring 1,000 human-curated questions across 162.9 hours of diverse, real-world audio to rigorously evaluate the reasoning and acoustic robustness of large audio language models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to understand the world not just by reading books, but by listening to it. For a long time, scientists have been building "Large Audio Language Models" (LALMs)—super-smart computer brains designed to hear sounds, recognize voices, and answer questions about what they hear. Think of these models as digital detectives that can listen to a conversation and tell you who is speaking, what they are saying, and even how they are feeling. But here is the catch: most of these detectives have only been trained in quiet, perfect classrooms where people speak clearly and slowly. They haven't really been tested in the messy, noisy, real world where people talk over each other, speak with different accents, or even struggle with speech due to medical conditions. If we want these AI detectives to be truly useful in real life, we need to see if they can handle the chaos of "the wild."
This is where a new project called ESCUCHA comes in. The name is Spanish for "Listen," and the researchers created it to be the ultimate test for these audio AI brains, specifically for the Spanish language. Instead of using clean, studio-recorded audio, they gathered 1,000 questions paired with 162.9 hours of real-world recordings. These clips range from a few seconds to over 80 minutes long and include everything from noisy street interviews to people with speech difficulties caused by conditions like ALS or stroke. The goal was to see if the AI could do more than just transcribe words; they wanted to see if the AI could actually reason about what it heard, like figuring out if a speaker lived longer than a predicted life expectancy based on a complex story.
The results of this "wild" test were a mix of impressive progress and stark reality checks. When the researchers pitted their best AI models against trained human experts, the humans won easily, scoring 90.10% correct answers while the best AI model, Qwen3-Omni-30B-A3B, only reached 74.40%. This gap suggests that while AI is getting good at listening, it still struggles with the deep, messy reasoning that humans do naturally. Interestingly, the study found that for many questions, a "cheat code" worked: if you simply typed the audio into a text transcript first and then asked a text-only AI to read it, that text-based system often performed better than the models that tried to listen directly. This implies that for a large chunk of these tests, the AI didn't need to understand the sound of the voice, just the words it contained.
However, the test also revealed a significant weakness. When the audio featured "non-normative" speech—meaning people with speech disorders or heavy accents—many of the AI models crashed, with some dropping their scores by 15 to 19 percentage points. This suggests that these models are still very fragile when faced with voices that don't sound "standard." The study concludes that while we are making strides, there is still a long way to go before AI can truly understand the full, messy spectrum of human speech, especially when it comes to the most vulnerable speakers. The ESCUCHA benchmark serves as a crucial map, showing us exactly where the AI is strong and where it is still lost in the noise.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.