← Latest papers
⚡ electrical engineering

MRMAD: A Multi-Round Multi-Audio Benchmark for Evaluating Acoustic Degradation Perception in Large Audio-Language Models

This paper introduces MRMAD, a novel multi-round, multi-audio benchmark designed to evaluate Large Audio-Language Models' ability to perceive, diagnose, and reason about acoustic degradations across speech, music, and sound, revealing that current models struggle with reliable degradation analysis despite their semantic understanding capabilities.

Original authors: Yize Li, Ningyuan Yang, Sile Yin, Sindhuja Thogarrati, Sung-En Chang, Andrew C. Singer, Xue Lin, Chuan-Che Huang, Shuo Zhang

Published 2026-08-25
📖 4 min read☕ Coffee break read

Original authors: Yize Li, Ningyuan Yang, Sile Yin, Sindhuja Thogarrati, Sung-En Chang, Andrew C. Singer, Xue Lin, Chuan-Che Huang, Shuo Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

We live in a world increasingly filled with machines that can hear. These systems, known as large audio-language models, are designed to listen to speech, music, and environmental sounds, then answer questions or follow instructions about what they hear. They have become quite good at understanding the meaning of a sentence or identifying that a sound is a dog barking. However, there is a fundamental aspect of hearing that these machines have largely ignored: the quality of the sound itself. In the real world, audio is rarely perfect. Recordings are often marred by background noise, echo, or the crackle of a poor internet connection. For a human listener, noticing these flaws is the first step in judging whether a recording is clear or broken. It remains an open question whether artificial intelligence can truly perceive these low-level imperfections or if it simply guesses based on the words it hears.

To investigate this, researchers from Northeastern University, Bose Corporation, and Stony Brook University created a new test called MRMAD. This benchmark is designed to see if audio-language models can actually hear the difference between a clean recording and a degraded one. The test covers three main areas of sound: human speech, music, and general environmental noises. It focuses on nine specific ways audio can be damaged, such as static from packet loss, the hollow sound of echo, or the muffled quality of low-pass filtering. The researchers did not just ask the models to listen once; they set up a conversation where the model hears a clean reference clip first, and then a degraded version, asking it to identify the type of damage, compare how severe it is, or rank multiple clips from least to most damaged. This multi-turn approach forces the model to hold the memory of the first sound in its mind while analyzing the second, mimicking how a human would compare two recordings.

The results of this extensive evaluation, which involved 8,400 questions and 18 different models, reveal a significant gap between current technology and human perception. Most of the open-source models tested performed only slightly better than random guessing. When asked to identify the specific type of damage or rank the severity of the distortion, these models often failed to distinguish between a clean sound and a broken one. Even the most advanced models, including some from major technology companies, struggled with the task. While one specific closed-source model from Google performed well, achieving high accuracy across the board, many other powerful systems faltered. This suggests that simply making a model larger or giving it more reasoning capabilities does not automatically grant it the ability to perceive audio quality. The study found that many models rely heavily on the semantic content—the words or the melody—rather than the acoustic texture of the sound.

The researchers dug deeper to understand why these models fail. They discovered that the models often do not use the clean reference audio provided in the first turn of the conversation. In many cases, removing the clean reference or replacing it with static noise did not significantly change the model's performance, indicating that the model was ignoring the comparison entirely. Furthermore, the internal representations of the audio within these models appear to smooth over the very details needed to detect damage. The digital tokens that represent the sound are often so similar between a clean clip and a degraded one that the model cannot tell them apart. This suggests that the problem lies not just in the reasoning step, but in how the sound is initially processed and preserved before the model even begins to think about it.

This work highlights a critical missing piece in the development of audio intelligence. While machines have become proficient at understanding what is being said or played, they have not yet mastered the ability to judge how well it is being said or played. The benchmark shows that current systems are not yet robust enough to handle the messy, imperfect audio conditions of the real world. Until these models can reliably detect and reason about acoustic degradation, their understanding of the auditory world will remain incomplete. The study concludes that building future systems that are truly robust will require a fundamental shift in how these models process sound, moving beyond high-level meaning to a deeper, more sensitive perception of audio quality.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →