CArtBench: Evaluating Vision-Language Models on Chinese Art Understanding, Interpretation, and Authenticity
This paper introduces CARTBENCH, a museum-grounded benchmark comprising four subtasks to evaluate Chinese art understanding, interpretation, and authenticity in vision-language models, revealing that while current models achieve moderate recognition accuracy, they significantly struggle with evidence-grounded reasoning, expert-style appreciation, and connoisseur-level authenticity discrimination.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a group of very smart, super-fast robots that can look at pictures and read text. We call them Vision-Language Models (VLMs). Right now, these robots are great at simple tasks, like saying, "That's a cat!" or "That's a red apple."
But what happens when you ask them to do something much harder? What if you put them in a museum and asked them to act like a curator or an art expert? Could they explain why a painting is beautiful? Could they tell the difference between a real masterpiece and a perfect fake?
This paper introduces CARTBENCH, a new "exam" designed specifically to test these robots on Chinese Art. Think of it as a specialized driving test for robots, but instead of driving cars, they are driving through the complex world of Chinese history, brushstrokes, and cultural meaning.
Here is a simple breakdown of what they did and what they found, using some everyday analogies:
1. The Problem: The "Tourist" vs. The "Expert"
Most current AI models are like tourists with a guidebook. They can point to a statue and say, "That's a dragon from the Tang Dynasty." That's good! But if you ask them, "Why does the brushwork here suggest it was painted by a specific master in the Ming Dynasty?" or "Is this a real painting or a very good copy?", they often get lost. They might guess confidently but be completely wrong.
The researchers wanted to know: Can these robots think like art experts, or are they just good at memorizing facts?
2. The Exam: CARTBENCH
To test this, the team built a benchmark (a test suite) called CARTBENCH. It's like a four-part obstacle course for the robots:
Part 1: The Detective (CURATORQA)
- The Task: The robot looks at a painting and answers questions. But here's the catch: it can't just guess. It has to point to the specific part of the image that proves its answer.
- The Metaphor: Imagine a detective looking at a crime scene photo. They can't just say "The butler did it." They have to say, "The butler did it because you can see mud on his shoes that matches the garden."
- The Result: The robots were okay at simple questions but terrible at connecting visual clues to historical facts (like guessing the time period based on the painting style).
Part 2: The Poet (CATALOGCAPTION)
- The Task: The robot has to write a long, structured appreciation of the art, divided into four sections: Background, Content, Artistic Features, and Overall Verdict.
- The Metaphor: It's like asking a robot to write a professional museum guidebook entry. It needs to sound like a human expert, not a Wikipedia summary.
- The Result: The robots could follow the format (they wrote four sections), but the content was shallow. They sounded like a student trying to sound smart, missing the deep cultural nuance and emotional depth of a real expert.
Part 3: The Creative Thinker (REINTERPRET)
- The Task: The robot has to give a fresh, creative interpretation of a famous painting, but it must stay grounded in what is actually visible.
- The Metaphor: Imagine a teacher asking a student to write a new ending to a classic story, but the student can't invent new characters that don't exist in the book.
- The Result: The robots struggled to be truly creative without making things up. They often got stuck in "safe" answers or hallucinated facts.
Part 4: The Forger Hunter (CONNOISSEURPAIRS)
- The Task: The robot is shown two paintings side-by-side. One is real, one is a fake. They look almost identical. The robot must pick the real one.
- The Metaphor: This is the "tough guy" test. It's like asking someone to spot a counterfeit $100 bill when the fake one is printed so well it looks perfect to the naked eye.
- The Result: The robots performed no better than random guessing (like flipping a coin). They couldn't spot the subtle differences in brushstrokes or ink flow that a human expert would catch.
3. The Big Reveal
The study tested 9 different AI models (including big names like Qwen, GPT, and Gemini). Here is what they found:
- The "Smart" Illusion: Some models got high scores on the easy questions. It looked like they were geniuses.
- The Reality Check: When the questions got harder—requiring them to link a visual style to a specific historical era or spot a fake—their performance crashed.
- The Language Gap: Models that were specifically trained to be bilingual (Chinese and English) did slightly better than general English-only models, but they still weren't "experts."
- The Authenticity Wall: The robots are currently terrible at spotting fakes. They rely on surface details (like "this looks detailed") rather than deep understanding (like "the brush pressure here is inconsistent with the artist's style").
4. Why Does This Matter?
If we start using AI to help manage museums, restore art, or even authenticate paintings, we need to know their limits.
- Risk: If a robot confidently tells a museum that a fake painting is real, the museum could lose millions of dollars.
- Goal: This paper isn't saying "AI is bad." It's saying, "AI is a great assistant, but it is not a replacement for a human art expert yet." We need to teach these robots to look deeper, not just at the surface.
In a Nutshell
CARTBENCH is a reality check for AI. It showed that while robots are great at naming things, they are still very bad at understanding the soul, history, and subtle details of Chinese art. They are like tourists who can read the map but haven't learned how to navigate the terrain. To get there, we need to teach them to "see" like experts, not just like cameras.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.