Do Vision Encoders Exhibit Human-like Color Thresholds?
This large-scale study reveals that over 50 pretrained vision encoders, including both convolutional networks and transformers, generally fail to exhibit human-like color discrimination thresholds, with self-supervised models performing best yet still showing weak alignment with human perceptual sensitivity.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Color Blindness of Our Smartest Machines
Imagine you are looking at a sunset. To your eyes, the sky isn't just one big blob of orange; it's a smooth gradient where some tiny shifts in color are easy to spot, while others blend together so perfectly you can't tell them apart. This is how human vision works: our brains are incredibly sensitive to some colors and less sensitive to others, creating a "map" of the world that isn't perfectly uniform. Scientists have known about this uneven sensitivity for decades, mapping out exactly how much a color needs to change before we notice it.
Now, imagine teaching a computer to see. We've built massive digital brains called "vision encoders" that can recognize cats, cars, and clouds by studying millions of photos. These machines are the stars of modern artificial intelligence. But here is the big question: Do these digital brains see color the way we do? Do they have their own version of that uneven "map," noticing the same subtle shifts in a red apple that a human would, or do they see the world through a completely different, perhaps much stranger, lens? This is the mystery a new study set out to solve, testing whether our most advanced AI has accidentally learned to see like a human, or if it's still colorblind in ways we never expected.
The Great Color Test: Do AI Brains See Like Us?
In this study, a team of researchers decided to put over 50 different AI vision models to the test. Think of these models as a huge class of students, ranging from older, classic "convolutional" networks to the newest, super-complex "transformers" and models that learned to see by reading text descriptions of images. The researchers wanted to know: Do these machines have human-like color thresholds?
To find out, they didn't just ask the AI to name colors. Instead, they created a game of "spot the difference." They picked 18 specific reference colors (like a specific shade of blue or green) and surrounded each one with a cloud of slightly different colors. For humans, there is a known "safe zone" around each color—a shape like a stretched-out oval (called a MacAdam ellipse)—where any color change inside that oval is invisible to the human eye. If you step outside that oval, even a tiny bit, you can suddenly see the difference.
The researchers then fed these colors into the AI models. They asked the AI: "Which of these surrounding colors looks most similar to the center one?" The AI answered by creating its own "similarity map" in its digital brain. If the AI was human-like, its map of "similar colors" should perfectly overlap with the human "safe zone" ovals.
The Results: A Mismatch in the Rainbow
The findings were a bit of a shocker. The study found that, generally, these AI models do not see color like humans do.
Even the best-performing models only matched human color sensitivity about 25% of the time (specifically, a score called mIoU was less than 0.25). To put that in perspective, if you were trying to match a human's color perception perfectly, these AI models were getting less than a quarter of the details right. They weren't completely blind to color, but their "map" of the rainbow was distorted.
Here is what the study discovered about the different types of AI students:
- The Self-Taught Students Won: Models that taught themselves by looking at pictures without any labels (self-supervised) did the best job. They were the closest to human vision, though still far from perfect.
- The Text-Learners Were Wildcards: Models that learned by matching images to text (like CLIP) were the most unpredictable. Some of them were at the very top of the class, while others were at the very bottom. They were polarized, meaning they either got it somewhat right or got it very wrong.
- Bigger Isn't Better: You might think a giant, super-complex AI model would see better than a smaller, older one. But the study showed that scaling up the model size didn't automatically fix the problem. In fact, some older, smaller models (like the classic VGG16) performed just as well as the massive, modern foundation models. This suggests that just making the brain bigger doesn't teach it to see color like a human.
- The Blue Problem: The AI struggled the most with blue colors. Just like humans find it hard to distinguish between very similar shades of blue, the AI had the hardest time here, often getting the "safe zones" completely wrong.
Why This Matters
The researchers suggest that human-like color sensitivity does not naturally appear just because an AI is trained on huge amounts of photos. The current way we train these models seems to prioritize other things (like recognizing objects or understanding text) over the fine, subtle details of color perception.
The study concludes that while these AI models have learned to organize colors in a structured way (they aren't just random noise), they organize it differently than we do. If you were to use these models for tasks that require perfect color matching—like restoring old paintings, checking if a fruit is ripe, or designing digital art—you might run into problems. The AI might think two colors are identical when a human can clearly see they are different, or vice versa.
In short, our digital eyes are getting better at seeing the "big picture," but they are still learning how to see the tiny, colorful details that make our world feel real to us. The gap between how a machine sees a sunset and how we see it is still wide open.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.