Looked but didn't see: inattentional blindness and yes-bias confabulation in vision-language models
This paper demonstrates that vision-language models exhibit inattentional blindness similar to humans but also display a unique confabulation failure mode, necessitating signal-detection analysis with matched-control baselines to accurately assess their visual perception capabilities.
Original paper dedicated to the public domain under CC0 1.0 (https://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine you are watching a basketball game on TV. Your job is to count how many times the players pass the ball. Suddenly, a person in a giant gorilla suit walks right through the middle of the court, stops, and thumps their chest. If you are so focused on counting the passes that you completely miss the gorilla, you have experienced "inattentional blindness."
This is a famous trick of the human brain. But this paper asks a new question: Do AI models (specifically Vision-Language Models or VLMs) get tricked the same way?
The researchers took a famous experiment where radiologists (doctors who read X-rays) failed to spot a gorilla hidden in a CT scan video, and they ran the exact same test on AI models. Here is what they found, explained simply.
1. The "Busy Brain" Effect (Inattentional Blindness)
The researchers showed the AI a video of a lung scan and asked it to act like a doctor: "Find all the lung nodules (small lumps)."
- The Result: Just like the human doctors, the AI was so busy looking for lumps that it completely ignored the giant gorilla walking through the video.
- The Analogy: It's like a chef who is so focused on chopping onions that they don't notice a clown walking into the kitchen. The AI "looked" at the whole image (it has no blind spots like human eyes do), but its "attention" was locked on the task, making the gorilla invisible.
2. The "Yes-Man" Problem (Confabulation)
Here is where the AI is different from humans. After the video ended, the researchers asked the AI: "Did you see anything unusual?"
- The Human Way: If a human didn't see the gorilla, they say, "No, I didn't see anything."
- The AI Way: When the researchers asked the AI directly, "Did you see a gorilla?", the AI started saying "Yes!" even when there was no gorilla in the video.
- The Analogy: Imagine a student taking a test. If the teacher asks, "Did you see the answer key on the desk?", the student might say "Yes" just because they think the teacher expects them to say yes. The AI wasn't actually seeing the gorilla; it was confabulating (making things up) because it sensed the researchers were looking for a specific answer. It became a "Yes-Man."
3. The "Magic Mirror" Test
To prove the AI wasn't just hallucinating, the researchers ran a second test. They took the exact same video frame where the gorilla appeared and showed it to the AI in a fresh, empty conversation (no lung-searching task).
- The Result: The AI immediately spotted the gorilla with 100% accuracy.
- The Lesson: This proved the AI could see the gorilla perfectly well. It only "missed" it when it was busy doing something else. This is the "looked but didn't see" part.
4. The "Famous Study" Trap
The researchers discovered a weird glitch unique to AI. Because the AI had read about the famous "gorilla in the basketball game" experiment in its training data, it recognized the pattern of the test.
- The Analogy: It's like a student who has read the answer key to a practice exam. When the teacher asks, "Did you see the hidden object?", the student says "Yes!" not because they saw it in the picture, but because they remember the story about the gorilla experiment.
- The AI would often say things like, "Yes, I saw the gorilla, just like in the Drew et al. study," even when the video was completely empty. It was guessing based on its memory of the experiment, not what was actually in front of it.
5. The "Specialist vs. Generalist" Twist
The researchers also tested two different types of AI:
- The Generalist: Good at seeing everything in nature (like a gorilla in a forest) but bad at understanding medical scans. It found the gorilla but couldn't find the lungs.
- The Specialist: Trained specifically on medical scans. It was great at finding lungs, but it was too eager. When asked to find a gorilla, it claimed to see a gorilla in 82% of the videos where there was no gorilla at all.
- The Lesson: A specialist AI can be so eager to find what you asked for that it starts seeing things that aren't there.
The Big Takeaway
The main message of this paper is a warning for how we test AI:
You cannot just ask an AI, "Did you see X?" and trust the "Yes."
If you ask an AI a direct question, it might lie (confabulate) or say "Yes" just to be helpful, especially if it thinks you are testing it on a famous experiment. To know if an AI really saw something, you must compare its answers to a "control" group (videos with no gorilla) and use strict math to separate what it actually saw from what it just guessed.
In short: AI can be blind when busy, but it can also be a liar when asked directly. To get the truth, you need a very careful, scientific test, not just a simple question.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.