← Latest papers
💻 computer science

Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical Scores

This paper introduces MSU-Bench, a human-curated benchmark comprising 1,800 generative question-answer pairs across textual and visual modalities to evaluate and improve the ability of Large Language and Vision-Language Models to understand complete musical scores through integrated reasoning over pitch, rhythm, harmony, and structure.

Original authors: Congren Dai, Yue Yang, Krinos Li, Huichi Zhou, Shijie Liang, Bo Zhang, Enyang Liu, Ge Jin, Hongran An, Haosen Zhang, Peiyuan Jing, Kinhei Lee, Z henxuan Zhang, Xiaobing Li, Maosong Sun

Published 2026-04-24
📖 5 min read🧠 Deep dive

Original authors: Congren Dai, Yue Yang, Krinos Li, Huichi Zhou, Shijie Liang, Bo Zhang, Enyang Liu, Ge Jin, Hongran An, Haosen Zhang, Peiyuan Jing, Kinhei Lee, Z henxuan Zhang, Xiaobing Li, Maosong Sun

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a brilliant, super-smart robot librarian who has read millions of books and can answer almost any question you ask. Now, imagine handing this robot a complete, multi-page musical score (like the sheet music for a symphony) and asking it to act like a music professor.

That is exactly what this paper, MSU-Bench, is about. The researchers wanted to see if today's most advanced AI models (like the ones powering ChatGPT or Google's Gemini) can actually "read" and understand a full piece of sheet music, not just a single note or a short snippet.

Here is the breakdown of their findings using some everyday analogies:

1. The Problem: The "Blind" Robot

The researchers found that while these AI models are great at reading text, they are terrible at reading visual sheet music (PDFs).

  • The "Lost in the Pages" Problem: If you ask a human, "What note is in the 7th bar?" they can count the lines and find it. But the AI often gets lost. It might look at bar 7 and accidentally describe bar 8, or it might just make up an answer because it's guessing.
  • The "Hallucination" Problem: This is like a student who didn't study for the test but tries to bluff their way through. The AI sees a question about a specific part of the music and, instead of admitting it can't see it clearly, it invents a fake answer. It says, "Oh, there's a loud drum there!" when there is actually silence.

2. The Solution: The "Textbook" Translation

To fix this, the researchers tried a clever trick. Instead of showing the AI the picture of the sheet music (PDF), they gave them a text-based translation called ABC Notation.

  • The Analogy: Imagine the sheet music is a complex, hand-drawn map. Reading the map is hard for a robot because the lines are squiggly and the symbols are tiny. ABC Notation is like taking that map and turning it into a simple list of GPS coordinates and street names (e.g., "Turn left at C, go straight for 4 beats").
  • The Result: When the AI read this "text list," it got much smarter. It could finally answer questions about chords, rhythm, and structure accurately. It proved that the AI knows music theory, but it just struggles to "see" the visual symbols on the page.

3. The Test: The "Four-Level Exam"

The researchers created a massive exam called MSU-Bench with 1,800 questions based on famous pieces by composers like Bach, Beethoven, and Chopin. They organized the test into four levels of difficulty, like climbing a ladder:

  • Level 1 (The Cover): "Who wrote this? What is the title? How fast should we play?" (Easy stuff).
  • Level 2 (The Details): "In the 5th measure, is there a sharp sign? What is the lowest note?" (Zooming in).
  • Level 3 (The Harmony): "What chord is playing here? Is it a happy major chord or a sad minor chord?" (Understanding the relationships between notes).
  • Level 4 (The Big Picture): "Where does the main melody theme appear? How does the song change from start to finish?" (Understanding the whole story).

4. The Results: A Mixed Report Card

When they tested over 15 different AI models, the results were surprising:

  • The Visual Gap: When looking at the actual sheet music images (PDFs), the AI models scored very low (around 20%). They were like students trying to read a book in a language they don't know.
  • The Text Boost: When given the text version (ABC), the scores jumped significantly (up to nearly 50%). The models could finally use their "brain" to solve the logic puzzles.
  • The "All-or-Nothing" Struggle: The hardest part wasn't just answering one question; it was answering all of them correctly in a row. The researchers used a metric called Level-wise Success Rate. It's like a video game: if you fail Level 1, you can't even try Level 2. Most AIs failed to keep their streak going past the first few levels.

5. The Future: Training the Robot

The researchers also tried "teaching" the AI by showing it examples (Fine-tuning).

  • The Good News: After a little bit of training, the AI got much better at both reading the text and understanding the images. It didn't forget how to do other things (like writing essays or solving math problems); it just added music to its skillset.
  • The Bad News: Even with training, the AI still struggles with the visual part. It's like teaching a person to read music, but they still have to squint at the page. They need better "eyes" (visual processing) to truly master sheet music.

The Bottom Line

This paper is a wake-up call for the AI world. AI is smart enough to understand music theory, but it's currently "blind" to the visual language of sheet music.

The researchers built this benchmark (MSU-Bench) to be a standard ruler for the future. They want to push AI developers to build models that can look at a piece of sheet music, find the right page, count the bars, and explain the music just like a human conductor would. Until then, if you want an AI to analyze a symphony, you might still need to give it a text translation first!

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →