ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval
The paper introduces ReasonAudio, the first benchmark designed to evaluate text-audio retrieval models on complex reasoning tasks like negation and duration, revealing that current state-of-the-art models and multimodal large language models significantly struggle with these challenges despite their proficiency in basic semantic matching.
Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to find a specific song in a massive library of sound effects, but instead of just humming a tune, you have to give the librarian a very specific set of logical instructions.
This paper introduces ReasonAudio, a new "test" designed to see how well computers can follow those tricky instructions when searching for sounds.
The Problem: Computers Are Bad at "Logic"
Right now, computers are great at matching sounds to words based on general vibes. If you ask for "a dog barking," a computer can usually find a clip with a dog.
But real life is messier. Imagine asking for:
- "A dog barking, but not a cat meowing." (Negation)
- "First a door slams, then a car honks." (Order)
- "A siren and a fire alarm happening at the exact same time." (Overlap)
- "A drum beat that lasts exactly 5 seconds." (Duration)
The authors found that current computers are terrible at these "logic puzzles." They get confused by the "but not," the "then," and the "how long." They tend to just grab anything that sounds like the words, ignoring the rules.
The Solution: A New "Exam" for Computers
To prove this, the team built ReasonAudio, a giant practice exam.
- The Library: They created 10,000 custom sound clips by mixing simple sounds (like a bell, a car, or a bird) together in specific patterns.
- The Questions: They wrote 1,000 questions that require the computer to use logic, not just memory.
- The Rules: The answers are 100% correct because they were generated by a computer program, so there's no human error in the grading.
The Results: Everyone Failed the Test
The researchers put 10 of the smartest, most advanced audio-search computers through this exam. The results were shocking:
- The "Two-Step" Students: Some computers try to first "listen" to the sound, write a description of it, and then search for that description. This was the worst approach. It's like trying to find a book by first asking a friend to describe the plot, then searching for the description. The friend's description was often too wordy or missed the point, so the computer got lost.
- The "Direct" Students: Other computers try to understand the sound and text directly. They did slightly better but still scored very low (around 8% accuracy).
- The "Super-Students" (MLLMs): The most advanced models (based on huge AI brains) did the best, but they still only got about 20% of the answers right. That's barely better than guessing.
The Big Surprise: Even though these "Super-Students" are famous for being great at logic when reading text or looking at pictures, they forgot how to use that logic when listening to audio. When they were fine-tuned to search for sounds, they seemed to lose their ability to understand "not," "before," or "how long."
Why Did They Fail?
The authors looked inside the "brain" of these computers and found two main issues:
- They ignore the rules: When a computer sees the word "dog," it grabs every clip with a dog, even if the instruction said "no dogs." It's like a librarian who hears "dog" and ignores the part of your request that said "no cats."
- They are messy: When the computer tries to match a text question to a sound, it doesn't organize them neatly. In the computer's mind, the "wrong" answer sometimes looks closer to the question than the "right" answer.
The Bottom Line
This paper doesn't say we can't search for audio yet. It just says that if you want a computer to find a sound based on complex, logical rules (like "the sound of rain, but only if it's not thunder, and it must be short"), current technology is not ready. The computers are currently "matching" words to sounds, but they aren't truly "reasoning" about them.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.