Investigating Modality Contribution in Audio LLMs for Music
This paper investigates the modality contributions in Audio LLMs for music by adapting the MM-SHAP framework to reveal that while higher-performing models rely heavily on textual reasoning, they still effectively utilize audio to localize key sound events.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart robot friend who loves music. You can talk to it, and it can listen to songs. You might ask, "What unusual sound starts this song?" and it should listen to the audio and tell you, "It's a doorbell!"
But here's the mystery: Is the robot actually listening, or is it just guessing based on the words you wrote?
This paper investigates exactly that. The researchers wanted to know if these "Audio Large Language Models" (Audio LLMs) are truly using their ears or if they are just relying on their brains (text processing) to figure out the answer.
The Detective Tool: MM-SHAP
To solve this mystery, the researchers used a special detective tool called MM-SHAP. Think of this tool like a "volume knob" for the robot's senses.
- How it works: The tool takes a question and a song, then it plays a game of "hide and seek." It randomly mutes (masks) parts of the text and parts of the audio.
- The Goal: It watches how much the robot's answer changes when a piece of the audio is silenced versus when a piece of the text is silenced.
- The Result: It gives a score showing how much the robot relied on the Audio versus the Text to make its final decision.
The Experiment: Two Robots, One Test
The researchers tested two different music robots (named Qwen-Audio and MU-LLaMA) using a quiz called MuChoMusic. This quiz has multiple-choice questions about songs, like "What instrument is playing?" or "What sound effect is heard?"
Here is what they found:
- The "Smart" Robot (Qwen-Audio): This robot got the most questions right. However, the detective tool revealed a surprise: It barely listened to the music! It relied heavily on the text of the question to guess the answer. It was like a student who didn't read the book but guessed the right answer because they knew the test format.
- The "Balanced" Robot (MU-LLaMA): This robot got fewer questions right, but it actually used the audio and text more equally. It was like a student who read the book and listened to the teacher, even if they still got some answers wrong.
The Big Takeaway: Being "accurate" (getting the right answer) didn't mean the robot was using the audio. In fact, the robot that used the audio the least was the one that got the most answers right. This suggests that for these specific multiple-choice questions, the robots are good at using logic and text clues, but they aren't necessarily "listening" in the way we hope.
Did They Ignore the Audio Completely?
Not exactly. The researchers dug deeper and looked at specific moments in the songs.
- The "Bell" Example: In one question asking about a "bell sound," the robot's internal logic did light up when the actual bell rang in the audio.
- The Catch: Even though the robot noticed the bell, it didn't seem to need that information to get the right answer. It could have guessed correctly just by reading the question options.
It's like a detective who sees a fingerprint at a crime scene but solves the case anyway because they already knew the suspect's name. The fingerprint (the audio) was there and noticed, but it wasn't the main reason the case was solved.
Why Does This Matter?
The paper concludes that these models are currently "text-heavy." They are great at reasoning with words, but they haven't fully learned to "listen" to music to solve problems yet.
The researchers also noted that the way these tests are designed (multiple-choice questions) might be too easy for the robots. They can often eliminate the wrong answers just by reading the text, without needing to hear the music at all.
In short: These music robots are very good at talking and thinking, but they are still learning how to truly listen. The tool developed in this paper helps us see exactly how much they are using their ears versus their brains.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.