← Latest papers
💬 NLP

BASS: Benchmarking Audio LMs for Musical Structure and Semantic Reasoning

This paper introduces BASS, a comprehensive benchmark comprising 2,658 questions across 12 tasks to evaluate the musical structure and semantic reasoning capabilities of audio language models, revealing that while current state-of-the-art models excel at lyric transcription, they significantly struggle with higher-level tasks like structural segmentation and artist collaboration.

Original authors: Min Jang, Orevaoghene Ahia, Nazif Tamer, Sachin Kumar, Yulia Tsvetkov, Noah A. Smith

Published 2026-02-05
📖 5 min read🧠 Deep dive

Original authors: Min Jang, Orevaoghene Ahia, Nazif Tamer, Sachin Kumar, Yulia Tsvetkov, Noah A. Smith

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a group of very smart robots that can listen to music and talk about it. You might think, "Great! They can probably tell me what's happening in a song, who is singing, and when the chorus starts."

The paper BASS (Benchmarking Audio LMs for Musical Structure and Semantic Reasoning) is essentially a giant, very difficult "final exam" designed to test exactly how good these music-listening robots really are.

Here is the breakdown of the exam, the students, and the results, using simple analogies.

1. The Exam: What is BASS?

Think of BASS not as a single test, but as a multi-room obstacle course with 12 different challenges. The researchers built this course using over 2,600 questions and nearly 140 hours of music from all kinds of genres.

The course is divided into four main "zones":

  • Zone 1: The Map Maker (Structural Segmentation)
    • The Task: Listen to a song and draw a map showing exactly when the Intro starts, when the Verse begins, and when the Chorus hits.
    • The Analogy: It's like watching a movie and trying to press a button every time the scene changes from "Kitchen" to "Car" to "Forest." The robots often get confused and think the "Forest" scene is actually the "Car" scene.
  • Zone 2: The Scribe (Lyric Transcription)
    • The Task: Listen to the song and write down the lyrics, but you must also know which part of the song you are writing for.
    • The Analogy: It's like a court stenographer who has to not only write down what the judge says but also label whether it's the "Opening Statement" or the "Closing Argument."
  • Zone 3: The Music Critic (Musicological Analysis)
    • The Task: Identify specific "genes" or traits in the music. Is the song sad? Is there a lot of drumming? Is the guitar distorted?
    • The Analogy: Imagine a sommelier tasting wine. Instead of saying "It's red," they have to say, "This has notes of oak, a hint of blackberry, and high tannins." The robots struggle to taste the subtle flavors of the music.
  • Zone 4: The Detective (Artist Collaboration)
    • The Task: In a song with two or more singers, figure out who is singing, when they start, when they stop, and whether they are rapping or singing.
    • The Analogy: It's like being at a crowded party and trying to track exactly which person is talking, when they stop, and what their voice sounds like, all while the music is playing. This was the hardest part of the exam.

2. The Students: Who took the test?

The researchers invited 14 different AI models to take this exam. These included the newest, most powerful "frontier" models (like Gemini 2.5 Pro) and several open-source models (like Qwen3-Omni).

Think of these models as students who have read millions of books and listened to millions of hours of audio. They are supposed to be geniuses.

3. The Results: How did they do?

The results were a bit of a shock. Even the "smartest" students didn't pass with flying colors.

  • The Overall Score: The best student (Gemini 2.5 Pro) only got about 26% of the questions right on average. Most other students scored even lower.
  • The Best Subject: The robots were surprisingly good at Lyric Transcription (writing down the words). This is like a student who is great at reading a script but terrible at understanding the plot. They rely heavily on the "language" part of their brain.
  • The Worst Subject: They struggled the most with Artist Collaboration and Structural Segmentation. They couldn't figure out who was singing when, or where the song sections began and ended.
  • The "Thinking" Factor: The researchers noticed that when they told the robots to "think step-by-step" before answering (like a human pausing to solve a math problem), the robots got better at some hard tasks, like figuring out who the artists were. However, for some tasks, thinking too much actually made them slower and more confused.

4. The Surprising Discoveries

The paper found a few interesting quirks about how these robots "think":

  • The "Cheat Sheet" Problem: When the researchers gave the robots only the song title and artist name (no audio), the robots actually got better at writing the lyrics!
    • What this means: The robots aren't really "hearing" the song to write the lyrics; they are just remembering the lyrics from their training data because they've seen the song title before. They are relying on memory, not listening skills.
  • The "Muddy Water" Problem: The researchers tried to help the robots by cleaning up the audio (removing the instruments and leaving only the vocals). They thought this would make it easier.
    • What this means: It actually made it harder. The robots needed the full, messy mix of instruments and vocals to understand the structure of the song. Taking away the instruments confused them.
  • The "Genre" Bias: The robots were much better at understanding Country and Rock music than Pop or Hip Hop.
    • What this means: The robots were likely trained mostly on Country and Rock data, so they are like a student who only studied for one specific type of test.

The Bottom Line

The paper concludes that while these AI models are getting better at understanding language, they are still very bad at understanding music as music. They can read the lyrics, but they struggle to understand the structure, the timing, and the complex interactions between different musicians.

The BASS benchmark is like a mirror showing the AI community exactly where their "musical ears" are still broken, so they can build better models in the future.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →