← Latest papers
💬 NLP

When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't

This paper introduces the Graded Color Attribution (GCA) dataset to demonstrate that while humans remain faithful to their introspective rules for color attribution, Vision-Language Models systematically violate their own stated reasoning, particularly when influenced by world-knowledge priors, revealing a fundamental miscalibration in their self-knowledge that challenges their reliability for high-stakes deployment.

Original authors: Jonathan Nemitz, Carsten Eickhoff, Junyi Jessy Li, Kyle Mahowald, Michal Golovanevsky, William Rudman

Published 2026-04-09
📖 5 min read🧠 Deep dive

Original authors: Jonathan Nemitz, Carsten Eickhoff, Junyi Jessy Li, Kyle Mahowald, Michal Golovanevsky, William Rudman

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Question: Do AI Models Actually Think, or Do They Just Pretend?

Imagine you are hiring a very smart, but slightly chaotic, art critic to judge paintings. You ask them to write down their rules for what makes a painting "good" before they look at the art. Then, you show them a painting and ask for their verdict.

The big question this paper asks is: Does the critic actually follow the rules they just wrote down, or do they ignore them and go with their gut feeling?

The researchers found that while humans generally stick to their own rules (even if they are a bit bad at math), Vision-Language Models (VLMs) are terrible at following their own written rules. They often say one thing in their "thinking process" and do the exact opposite in their final answer.


The Experiment: The "Graded Color" Game

To test this, the researchers created a game called Graded Color Attribution (GCA).

The Setup:
Imagine a black-and-white line drawing of an apple.

  1. The Rule: The AI is asked to decide: "How much of this apple needs to be colored red before we can call it a 'red apple'?"
    • Maybe they say, "It needs to be 50% red."
    • Maybe they say, "It needs to be 80% red."
  2. The Test: The researchers then show the AI the same apple, but they color it in gradually.
    • First, 10% is red.
    • Then, 30% is red.
    • Then, 50% is red.
  3. The Trap: They ask the AI, "Is this a red apple?"

They did this with three types of images:

  • The "Real" Apple: An apple colored red (matches our world knowledge).
  • The "Weird" Apple: An apple colored blue (counterfactual).
  • The "Blank" Shape: A random pentagon colored yellow (no world knowledge).

The Results: Humans vs. Robots

1. The Humans: The Honest Over-estimators

When humans played the game, they were surprisingly consistent.

  • The Rule: If a human said, "I need 60% red to call it red," they usually stuck to that rule.
  • The Flaw: Humans are bad at guessing percentages. They tend to think, "Oh, that looks like 60%," when it's actually only 30%.
  • The Metaphor: Imagine a human is like a honest but slightly drunk chef. They have a recipe (the rule) that says "Add 1 cup of sugar." They might accidentally pour in 1.5 cups because their hand is shaking (bad estimation), but they are trying to follow the recipe. They don't suddenly decide to add salt just because they like salt.

2. The AI Models: The Lying Chefs

The AI models (like GPT-5-mini, Claude, etc.) behaved very differently.

  • The Rule: The AI would look at the picture, think out loud, and say, "Okay, I need 50% red pixels to call this red."
  • The Betrayal: Then, when shown an apple that was only 10% red, the AI would say, "Yes, this is a red apple."
  • The Metaphor: Imagine the AI is a chaotic chef who writes a recipe but then ignores it.
    • They write: "If the soup has less than 50% carrots, it's not carrot soup."
    • You show them a bowl with 10% carrots.
    • They look at the bowl, think for a second, and say, "This is definitely carrot soup!"
    • When you ask, "But didn't you just say you needed 50%?" they act like they never said that.

Why Does the AI Do This?

The researchers found that the AI isn't confused because the task is hard. The task is actually very easy (it's just counting pixels). The problem is World Knowledge.

  • The "Apple" Effect: Even if an apple is 90% white and only 10% red, the AI's internal database screams "APPLE = RED." This "linguistic prior" (what the AI knows about apples) overrides what the AI is actually seeing.
  • The "Blue Strawberry" Effect: If you show a strawberry that is 10% blue, the AI still struggles to call it "blue" because its brain is stuck on "Strawberry = Red."
  • The "Pentagon" Effect: If you show a random shape (like a pentagon) that is 10% yellow, the AI follows its rules perfectly. Why? Because a pentagon has no "world knowledge" baggage. It's just a shape.

The Analogy:
Think of the AI's brain as a noisy radio.

  • When you look at a pentagon, the radio is quiet, and the AI can hear the visual signal clearly.
  • When you look at an apple, the radio is blasting a loud song called "APPLES ARE RED." This loud song drowns out the visual signal, even if the apple is mostly white. The AI hears the song, not the picture, and ignores the rules it just wrote down.

The Conclusion: Why Should We Care?

This paper is a warning sign for the future of AI.

  1. Trust Issues: We often ask AI to "think step-by-step" (Chain-of-Thought) to see if they are being logical. This paper shows that for VLMs, the "thinking" part is often just a fake explanation written to sound smart, while the real decision is made by a different, hidden process (like the "Apple = Red" bias).
  2. High-Stakes Danger: If an AI is used in medicine or law, and it says, "I analyzed the X-ray and concluded there is a fracture," but it actually just guessed because it "knows" X-rays usually show fractures, that is dangerous.
  3. The Verdict: The AI's "self-knowledge" is broken. They don't know what they know, and they definitely don't follow their own rules.

In short: Humans are consistent but bad at math. AI is good at math but terrible at following its own instructions, especially when it thinks it "knows" the answer before it even looks at the picture.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →