MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs
This paper introduces MuseBench, a comprehensive benchmark comprising over 4,000 expert-validated questions across diverse audiovisual arts, designed to evaluate the intent-level reasoning capabilities of multimodal large language models and revealing a significant performance gap between current AI systems and human experts in understanding creative artistic intent.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are watching a movie. A standard video AI might tell you, "There is a man crying in a dark room." It sees the pixels and recognizes the objects.
But MuseBench asks a much harder question: "Why did the director choose to make the room dark and the camera shake? Is it to make you feel trapped, or to show that the character is losing their mind?"
This paper introduces MuseBench, a new "exam" designed to test if AI can understand the artistic soul of videos, not just the surface details. Here is how it works, broken down simply:
1. The Problem: AI is a "Spoiler," Not a "Critic"
Current AI models are great at spotting things (like "a dog" or "a car"). But when it comes to art—movies, paintings, theater, and video games—they struggle to understand the intent.
- The Analogy: Imagine a student who can memorize the entire script of a play but doesn't understand why the actor paused for three seconds before speaking. They know what happened, but not why it was done that way.
- The Gap: Existing tests only ask, "What is happening?" MuseBench asks, "What was the artist trying to make you feel, and how did they do it?"
2. The Solution: The "Video Essay" Library
To build this test, the researchers didn't just grab random clips. They went to the internet (YouTube, Bilibili, TikTok) and collected over 10,000 "video essays."
- What is a video essay? Think of it as a movie critic talking over a clip, pointing out exactly why a specific lighting choice creates fear, or why a specific camera angle makes a character look powerful.
- The Process: The researchers used AI to turn these expert critiques into exam questions. They created 4,016 questions covering four main art worlds:
- Cinema (Movies/TV)
- Static Visual Arts (Paintings/Photos)
- Stage Performance (Theater/Dance)
- Game Arts (Video Games)
3. The Exam Format: It's Not Just "A, B, C, D"
Real art is often open to interpretation. A painting might be about "loneliness" AND "peace" at the same time.
- The Old Way: Most AI tests force you to pick one single right answer (like a multiple-choice quiz).
- The MuseBench Way: They use a mix of single-choice (pick the best answer) and multi-select (pick all the valid interpretations).
- Analogy: If you ask, "What is this song about?" a simple test might force you to pick "Sadness." MuseBench allows you to pick "Sadness" AND "Hope" if both are supported by the video. This tests if the AI can handle the complexity of human art.
4. The Results: The AI is Still a "Novice"
The researchers tested 28 of the smartest AI models available today (including big names like GPT-4, Claude, and Gemini) on this exam.
- The Score: Even the best AI only got about 48% correct.
- The Human Score: Human art experts got 87% correct.
- The Takeaway: The AI is basically guessing on half the questions. It hasn't learned the "language" of art yet.
5. Where Do They Fail?
The paper found some funny and specific patterns in how the AI fails:
- The "Game" Blind Spot: The AI was terrible at understanding video games. It seems the AI has read a lot of movie scripts but hasn't "played" enough games to understand how game design creates feelings.
- The "First Option" Bias: When the AI didn't know the answer, it had a weird habit of just picking the first option (Option A) more often than humans do.
- The "Salient" Trap: On questions with multiple correct answers, the AI usually found the most obvious answer but missed the subtle ones. It's like seeing the big tree but missing the forest.
Summary
MuseBench is a reality check for AI. It proves that while AI can describe a video, it is still very bad at understanding the creative choices behind it. To get better, AI needs to stop just "looking" at pixels and start learning the "vocabulary" of artists, directors, and game designers.
The Bottom Line: We have AI that can describe a painting, but we don't yet have AI that can truly appreciate it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.