Diagnosing Vision Language Models' Perception by Leveraging Human Methods for Color Vision Deficiencies
This paper evaluates Large-scale Vision-Language Models (LVLMs) using the Ishihara Test and finds that, despite possessing factual knowledge about color vision deficiencies, the models fail to simulate the altered perceptual experiences of affected individuals, instead defaulting to normative color perception and highlighting a critical gap in their ability to support inclusive accessibility.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, well-read robot assistant. You've taught it everything about the human eye, color blindness, and how people with different types of vision see the world. It can write a perfect essay explaining what it's like to be colorblind.
But then, you show it a picture and ask, "If you were colorblind, what number would you see in this picture?"
Instead of seeing the world through the eyes of a colorblind person, the robot just looks at the picture with its own "perfect" eyes and tells you the answer a normal person would see. It knows the facts about color blindness, but it can't actually feel or simulate the experience.
That is the core discovery of this paper.
The "Ishihara" Test: The Robot's Eye Exam
To test this, the researchers used a famous tool called the Ishihara Test. You've probably seen these before: they are circles filled with colorful dots.
- For a person with normal vision: The dots form a clear number, like "12" or "8."
- For a person with red-green color blindness: The colors blend together, and they might see a completely different number (like "3") or nothing at all.
The researchers treated these pictures like a diagnostic tool for robots. They asked various large AI models (the "robots") to look at these pictures and pretend to be different types of colorblind people.
The Three Ways They Checked the Robots
The researchers didn't just ask, "What number do you see?" They looked at the robots in three different ways, like a doctor checking a patient:
The "What You Say" Check (Generation):
They asked the robot to simply say the number it saw.- The Result: Even when told, "You are red-green colorblind," the robot almost always gave the answer a person with perfect vision would give. It couldn't "hallucinate" the wrong number that a colorblind person would actually see. It was stuck in "normal vision mode."
The "How Sure Are You?" Check (Confidence):
They measured how confident the robot was in its answer.- The Result: The robot was very confident when answering for "normal vision." But when asked to answer for a colorblind person, it became very confused and unsure. It didn't just guess the wrong number; it didn't even know how to be unsure in the right way. It was like a student who knows the answer to a math problem but gets lost when asked to solve it backwards.
The "Inside the Brain" Check (Internal Representation):
They looked deep inside the robot's "brain" (its internal data layers) to see if it was actually processing the image differently based on the instructions.- The Result: Even deep inside, the robot wasn't changing its perspective. It was processing the image as if it had perfect vision, regardless of what it was told to pretend. It was like a person wearing a blindfold but still describing the room as if they could see it perfectly.
The "Doctor" and the "Pattern Learner"
The researchers wondered: "Maybe the robots just need more medical training?" or "Maybe they are just memorizing the dot patterns?"
- They tested a robot trained specifically on medical data. Result: It still failed to simulate color blindness.
- They trained a robot on thousands of fake dot patterns to see if it was just memorizing shapes. Result: It got better at seeing numbers when it had perfect vision, but it got worse at pretending to be colorblind.
The Big Takeaway
The paper concludes that while these AI models are great at talking about color blindness, they are terrible at experiencing it.
Think of it like this: You can read a book about what it feels like to be underwater, but reading the book doesn't mean you can hold your breath for five minutes. These AI models have read the "book" about color blindness, but they cannot actually "hold their breath" and see the world through a colorblind person's eyes.
This matters because as we start using these robots to help people navigate the world, read maps, or understand charts, we need to make sure they don't accidentally give instructions that only work for people with perfect vision, leaving others behind. The paper shows that right now, these robots are still "vision-normal" at their core, even when they try to pretend otherwise.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.