← Latest papers
⚡ electrical engineering

Robustness assessment of large audio language models in multiple-choice evaluation

This paper reveals that large audio language models exhibit significant sensitivity to choice ordering and paraphrasing in multiple-choice evaluations, prompting the authors to propose a more robust evaluation protocol and metric to address these variabilities across three benchmarks and four models.

Original authors: Fernando López, Santosh Kesiraju, Jordi Luque

Published 2026-06-25
📖 4 min read☕ Coffee break read

Original authors: Fernando López, Santosh Kesiraju, Jordi Luque

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are taking a multiple-choice test, but instead of reading a book, you are listening to a sound clip. You have to answer a question like, "Why did the man say 'wait'?" with options like "To answer the phone" or "To find his keys."

This paper is about a group of researchers who decided to check if the new "super-smart" computers (called Large Audio Language Models) that listen to these sounds are actually listening, or if they are just guessing based on the words on the page.

Here is the breakdown of their investigation using simple analogies:

1. The Problem: The "Magic Trick" of Test-Taking

The researchers found that these AI models are like students who are great at spotting patterns in the test paper itself, rather than understanding the actual lesson.

  • The "Order" Trick: If you shuffle the order of the answers (putting the correct one first instead of last), the AI's score changes dramatically. It's like a student who only knows the answer is "C" because it's always in the third spot, not because they know the material.
  • The "Word Choice" Trick: If you rewrite the question or the wrong answers using different words (paraphrasing), the AI gets confused. It turns out the AI was relying on specific "clues" in the text to guess the right answer without ever really processing the audio.

2. The Experiment: The "Torture Test"

To see how robust (strong) these AI models really are, the researchers didn't just run them once. They ran them through a "Torture Test" with three different audio datasets (MMAU, MMAR, MMSU) covering sounds like speech, music, and noise.

They took the same audio clip and the same question, but they created many different versions of the text:

  • Shuffling the deck: They rearranged the order of the four answer choices in every possible way (24 different orders!).
  • Rewriting the script: They used other AIs to rewrite the question and the answers in different ways, keeping the meaning the same but changing the words.
  • The "Mix": They combined all these changes at once.

3. The Shocking Results

The results showed that these "super-smarts" are actually quite fragile.

  • The "Text-Only" Ghost: When the researchers took away the audio entirely and just gave the text to a standard text-based AI, that text-only AI still got a surprisingly high score (48.3% on one test). This proves the test questions had "shortcuts" that allowed the AI to guess correctly without hearing a single sound.
  • The "Distractor" Trap: The biggest drop in performance happened when they rewrote the wrong answers (distractors). It seems the models were looking at the wrong answers to figure out the right one. If the wrong answers were rewritten, the models got lost.
  • The "Longer is Better" Bias: The researchers noticed the models had a weird habit: they loved picking the longest answer. Even if the longest answer was wrong, the AI picked it about 50% of the time. It's like a student thinking, "The teacher wouldn't write a short, simple answer for the hard question; it must be the long, complicated one."

4. The New Scorecard

Because the standard "Accuracy" score (just counting how many times they got it right) was misleading, the researchers proposed a new way to grade these models.

  • The "Consistency" Score: Instead of just asking, "Did you get it right?" they ask, "Did you get it right every single time, even when we changed the order or rewrote the words?"
  • They call this the Correctness Rate (CoR). If an AI gets the answer right 90% of the time normally, but only 20% of the time when the test is slightly tweaked, its CoR is low. This tells us the model isn't truly "smart"; it's just good at one specific version of the test.

5. The Winners and Losers

  • Audio Flamingo 3 was the overall champion, getting the highest scores and staying relatively steady even when the test was tweaked.
  • Qwen2.5-Omni was the most "stubborn" (robust) when it came to the tricky "wrong answer" rewrites.
  • Audio Flamingo 2 struggled the most, showing it was very sensitive to changes in the text.

The Bottom Line

The paper concludes that we can't just trust the standard test scores for these audio AIs. They are often "cheating" by reading the test paper instead of listening to the sound. To know if they are truly good, we need to shake up the test—change the order, rewrite the words, and see if they still get it right. If they fail the "shuffled" test, they aren't ready for the real world yet.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →