← Latest papers
💻 computer science

Whose Art Counts? Model- and Prompt-Dependent Associations in Vision-Language Judgments of Museum Art

This paper introduces an archive-conditioned audit protocol to evaluate how vision-language models and prompt structures interact with museum metadata when assessing artistic value, demonstrating through a Metropolitan Museum of Art case study that "influence" prompts yield the most consistent category-based differences while highlighting the limitations of cross-model and cross-prompt generalization.

Original authors: Manpreet Singh, Nandakishor Reddy Pulagam, Rhythm Bhatia, Rahul Joshi

Published 2026-08-14
📖 5 min read🧠 Deep dive

Original authors: Manpreet Singh, Nandakishor Reddy Pulagam, Rhythm Bhatia, Rahul Joshi

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you've just built a super-smart robot librarian who has read millions of books and seen billions of pictures. You ask it, "Which of these paintings is a masterpiece?" or "Which one changed history?" You expect the robot to look at the art and give a fair answer based on beauty and skill. But here's the catch: this robot doesn't just "see" art like a human does. It sees patterns. It learned by matching pictures to words on the internet. If the internet mostly said "great artists" were men, the robot might learn that "great art" looks like a man's work, even if it's wrong. This is the world of Vision-Language Models (VLMs): computers that connect images to words. The big question isn't just "Is the robot smart?" but "What kind of world did the robot learn from?" and "Does the robot's answer change if we ask the question differently?" Museums are full of these robots now, helping people find art, but if the robot is biased, it might hide the stories of women and other underrepresented creators, making them seem less important than they really are.

So, a team of researchers decided to put these robot librarians to a very specific test. They didn't just ask the robots to guess; they set up a strict game to see if the robots were playing fair. They grabbed 445 digital pictures of art from the Metropolitan Museum of Art. But here's the twist: they knew that museum records are messy. Many old artworks don't have a clear name attached to them. So, they split the group into two: a small group of 61 pieces where the creator's name was known (and they guessed the gender based on the name), and a huge group of 384 pieces where the creator was a mystery. They wanted to see if the robots gave "higher scores" to the named male artists compared to the named female artists.

But they didn't just ask one question. They asked the robots four different ways: "Is this a masterpiece?", "Is this quality work?", "Is this influential?", and a control question just to check if the robot was paying attention. They ran this test on four different robot brains (OpenAI CLIP, OpenCLIP B/32, OpenCLIP L/14, and SigLIP).

Here is what they found, and it's a bit of a plot twist: The robots didn't agree.

It wasn't a simple story where "all robots are sexist." Instead, the answer depended entirely on which robot you asked and how you asked.

  • When they asked about "Influence" (how much an artist changed history), two of the robots (OpenAI CLIP and OpenCLIP L/14) gave a clear signal: they rated the male-inferred works higher than the female-inferred ones. It was like the robot was saying, "Oh, history books say men changed things, so I'll give them a higher score."
  • However, when they asked about "Quality" or "Masterpiece," the results were a mess. One robot might say "Men are better," another might say "Women are better," and a third might say "I don't see a difference."
  • One robot, SigLIP, actually gave the opposite result, rating the female-inferred works higher, though the researchers noted the numbers were a bit wobbly and hard to pin down.

The most important thing the paper says is that you cannot just average these answers together. If you took the scores from all four robots and mashed them into one big number, you would get a lie. The robots are too different. One robot's "high score" is not the same as another robot's "high score." The researchers argue that the "bias" isn't a fixed flaw in the technology; it's a measurement problem. The result changes depending on the tool you use and the exact words you type.

The study also points out that the museum's own records played a huge role. Because 384 out of the 445 artworks had no known creator, the robots couldn't even compare them. The "named" artists were a tiny, specific slice of the museum's collection. The researchers warn that we can't say these robots hate women artists in general; we can only say that in this specific test, with these specific words and this specific set of pictures, some robots showed a pattern of favoring male names.

In the end, the paper doesn't say "Robots are broken" or "Robots are perfect." It says, "Be careful." If a museum uses a robot to rank art, the ranking might just be a reflection of how the robot was trained and how the question was asked, not the true value of the art. The researchers suggest that instead of looking for one "truth," we need to listen to the disagreement between the robots. If one robot says "This is a masterpiece" and another says "Nope," that disagreement is actually the most honest answer we can get: it tells us that the question is harder than we thought, and that the robot's opinion is just one noisy voice in a very loud room.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →