Systematic comparison of color representations between humans and deep neural networks: towards predicting human color perception in a vast color space
This study systematically compares color representations across self-supervised, supervised, and CLIP-trained deep neural networks using Gromov-Wasserstein Optimal Transport to demonstrate that while early layers of all paradigms align with human perception, only CLIP maintains this structural congruence at the output, enabling reliable predictions of human color perception across vast, previously unexplored color spaces.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine trying to map out how humans see the entire world of colors. For a long time, scientists have done this by asking people to compare pairs of colors, like "Is this red more like orange or purple?" But this is like trying to map a whole continent by only walking a few steps at a time. It takes forever, and we've only explored the "neighborhoods" of similar colors, leaving the vast, global landscape of 4,000+ colors a mystery.
To solve this, the researchers decided to use Deep Neural Networks (DNNs) as digital detectives. Think of these networks as artificial brains that have learned to see. The big question was: Which type of artificial brain sees colors the most like a human does?
They tested three different "training camps" for these digital brains:
- Self-Supervised Learning (SSL): The brain learns by looking at millions of photos on its own, trying to figure out patterns without any help.
- Supervised Learning (SL): The brain learns by looking at photos while a teacher points and says, "That's a cat," "That's a dog."
- CLIP: The brain learns by matching photos with their written descriptions (like "a red apple" or "a blue sky").
To see if these digital brains matched human vision, the researchers used a special mathematical tool called Gromov-Wasserstein Optimal Transport (GWOT). You can think of this as a super-precise "shape-shifter" that tries to fold the digital brain's map of colors and the human's map of colors together to see if they fit perfectly.
Here is what they found:
- The Early Layers: When they looked at the "early thinking" parts of all three types of networks, they all did a great job. They were like apprentices who had already learned to distinguish a red apple from a green one just as well as a human could.
- The Final Verdict: However, as the information traveled deeper into the networks, things changed.
- The Self-Supervised and Supervised networks started to drift away from human perception by the time they reached their final answer. They developed their own unique, non-human ways of organizing colors.
- CLIP was the winner. It was the only network that kept its "human-like" color map all the way to the very end. Even after processing all the information, its final view of color still looked structurally like how humans see it.
The Big Picture:
The researchers then used this winning network (CLIP) to explore a massive territory of 4,096 colors—a scale that would take humans decades to test in a lab. The network didn't just stop at the colors we know; it generated a complete, structured map of how these thousands of colors relate to one another.
In short, this paper shows that by training an AI to match images with words (CLIP), we get a tool that sees colors just like us. We can now use this tool to predict how humans would perceive vast, untested worlds of color, acting as a compass for future scientific exploration without needing to run endless, time-consuming experiments on real people.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.