← Latest papers
💬 NLP

EduArt: An educational-level benchmark for evaluating art history knowledge in large language models

This paper introduces EduArt, a psychometrically robust, educational-level benchmark comprising 871 human-authored questions in multiple languages and formats, which reveals that while large language models excel at multiple-choice art history recognition, their performance drastically drops on open-ended and error-identification tasks, highlighting a critical gap between recognition and the ability to reliably deploy art-historical knowledge.

Original authors: Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza

Published 2026-07-03
📖 4 min read☕ Coffee break read

Original authors: Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to test how smart a group of new students is at art history. For a long time, the standard test has been a multiple-choice quiz. You show them a picture of a painting and ask, "Who painted this?" with four options: A, B, C, or D.

The paper "EduArt" argues that this method is like testing a chef's cooking skills by only asking them to point at a picture of a burger and say, "Yes, that's a burger." It's too easy. The smartest AI models (the "students") are now getting almost 100% on these multiple-choice questions. They aren't necessarily understanding the art; they are just getting really good at guessing the right letter from a list.

The New Test: EduArt
The authors created a new, tougher exam called EduArt. Instead of just multiple-choice, they used real questions from Italian high school textbooks and US college entrance exams. They didn't make these questions up with a computer; real human teachers wrote them.

Think of EduArt as a "multi-skill obstacle course" for AI. Instead of just picking a letter, the AI has to:

  • Fill in the blanks: "The artist used [______] technique."
  • Spot the error: "Find the wrong word in this sentence about the painting."
  • Drag and drop: "Put the correct word in the right spot in the text."
  • Explain their work: "Why did you choose that answer?"

What They Found
The researchers tested 12 different AI models (like the brains of different companies) on this new course. Here is what happened, using some simple metaphors:

  1. The "Recognition" Trap: When the AI was just asked to pick an answer from a list (multiple-choice), the top models scored over 94%. They looked like art history geniuses. But when the test changed to asking them to write the answer or find the mistake, their scores crashed. One model that was 94% good at picking answers dropped to 6% when asked to find errors.

    • The Metaphor: It's like a student who can ace a trivia game show by buzzing in fast, but fails completely when asked to write an essay on the same topic. They know how to recognize the answer, but they don't know how to use the knowledge.
  2. The "Picture" Paradox: You might think that showing the AI the painting would make it easier. Surprisingly, the AI actually did worse when the image was included.

    • The Metaphor: It's like giving a student a map and a compass, but they get so distracted by looking at the map that they forget how to walk. The AI seemed to rely more on its text memory than on actually "seeing" the picture. When the picture was there, it got confused; when it was just text, it did better.
  3. The "Explain Yourself" Effect: The researchers asked the AI to write a short reason for its answer.

    • The Metaphor: Imagine a student who usually guesses the right answer on a test. When the teacher says, "Okay, now tell me why you picked that," some students freeze and get the answer wrong. Others get more confident and get it right.
    • The Result: It depended entirely on the "family" the AI came from. Some AI families (like Claude) actually got better when they had to explain themselves. Others (like Google's Gemini and OpenAI's GPT) got worse. It seems that forcing them to think out loud sometimes trips them up.

The Big Takeaway
The main point of the paper is that knowing the answer and being able to use the answer are two different skills.

If you only test AI with multiple-choice questions, you are getting a misleadingly high score. It's like judging a car's performance only by how well it can park in a straight line, ignoring whether it can drive up a mountain or turn a corner.

For art history experts who want to use AI to help analyze paintings or write descriptions, this paper warns: "Don't trust the multiple-choice scores." Just because an AI can pick the right name for a painting doesn't mean it can write a correct paragraph about it or spot a fake. To know what an AI can really do, you have to test it with the messy, open-ended tasks that real scholars actually face.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →