← Latest papers
💻 computer science

Human-AI Perceptual Alignment by Playing Hues and Cues

This paper introduces a novel evaluation framework using the board game Hues and Cues to assess human-AI perceptual alignment in color, revealing that while Contrastive Vision-Language Models replicate human biases for concrete objects, they systematically fail in abstract domains due to semantic misclassification and uncertainty collapse, a problem mitigated more effectively by curated pre-training datasets than by massive uncurated corpora.

Original authors: Nuria Alabau-Bosque, Jorge Vila-Tomás, Paula Daudén-Oliver, Pablo Hernández-Cámara, Valero Laparra, Jesús Malo

Published 2026-08-10
📖 6 min read🧠 Deep dive

Original authors: Nuria Alabau-Bosque, Jorge Vila-Tomás, Paula Daudén-Oliver, Pablo Hernández-Cámara, Valero Laparra, Jesús Malo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to see the world, not just by recognizing shapes, but by understanding the feeling of things. In the world of artificial intelligence, there is a special type of brain called a "Vision-Language Model." Think of these models as super-smart students who have read billions of books and looked at billions of photos from the internet. They are great at saying, "That is a dog," or "That is a car." But can they understand the color of a dog? Can they know that a banana is usually yellow, not because it's physically yellow in every single photo, but because our brains have a "memory" that bananas are yellow? This is the tricky part: human color isn't just about light hitting our eyes; it's about our culture, our memories, and our shared stories. Scientists want to know if these AI students are learning to see like humans, or if they are just guessing based on the weird patterns they found in their massive data piles.

To find out, a team of researchers decided to play a game. They didn't use boring charts or complex math tests. Instead, they used a board game called "Hues and Cues." Imagine a giant board with 480 tiny squares, each a slightly different color, arranged like a rainbow map. The game works like this: someone says a word (like "Banana" or "Hope"), and you have to point to the square on the board that best matches that word. It sounds simple, but it reveals how our brains connect words to colors. The researchers used this game to test 162 different AI models, asking them to play along with 325 real humans. They wanted to see if the AI's "guesses" matched the crowd's "memory," or if the AI was getting lost in its own head.

Here is what happened when the robots sat down to play.

The Setup: A Digital Board Game
The researchers built a digital version of the "Hues and Cues" board. This board had 480 distinct color squares, mapped out so they could be measured scientifically. They gathered a massive group of 325 people to play the game on their own phones and computers. Each person was asked to match 100 different words to the best color on the board. Some words were easy and physical, like "Banana" or "Blood." Others were tricky and abstract, like "Hope," "Feminism," or "Pop Culture."

To make sure the AI wasn't deviating from the protocol, the researchers gave the models the exact same digital board. The AI's job was to look at a word and pick the top 5 color squares it thought matched best. The researchers then compared the AI's choices to the "Human Consensus"—the average answer given by the real people. They even calculated a "Human Error Rate," which is basically how much humans disagree with each other. This gave them a perfect baseline: if an AI is as good as a human, it should make about the same amount of "mistakes" as a person does.

The Big Discovery: The "Blue" Crash
The results were a mix of impressive wins and a very strange, systematic failure.

First, the good news: When the words were about real, physical objects, the AI was surprisingly good. For concrete things like "Banana," "Pumpkin," or "Pig," the AI models often picked the exact same colors as the humans. They understood that bananas are yellow and pigs are pink. In fact, for some of these physical objects, the AI was so precise that it even matched the "idealized" memory color humans have, rather than just the literal color of a specific photo. This suggests that when there is a clear, physical reality, the AI can learn the human "vibe."

However, when the game moved to abstract or cultural concepts, the AI hit a wall. For words like "Hope," "Rage," "Feminism," or even "Tractor" (which has a specific cultural meaning in Spain), the AI stopped making sense. Instead of picking a color that matched the feeling or the culture, the models started doing something bizarre: they all collapsed into a single, default color.

The researchers found that when the AI didn't know the answer, it didn't guess randomly. Instead, it almost always picked a specific shade of blue (a coordinate they call "I18" on the board). It was as if the AI's brain short-circuited and said, "I don't know what this means, so I'll just pick blue." This happened over and over again, regardless of whether the word was "Winter," "Stitch" (from the movie), or "Axolotl."

Why Did This Happen?
The team dug deeper to find out why the AI was so obsessed with blue. They tested the models with "nonsense" words—things that have no meaning at all, like random letters or stop words like "the" and "and." They found that about 42% of the models had a built-in bias. When they saw nonsense, they didn't guess randomly; they defaulted to specific colors, often that same blue or sometimes yellow.

This revealed a critical flaw: the AI wasn't really "understanding" the word "Hope." It was just following a pattern it learned from its training data. Because the internet has so many pictures of the sky and water (which are blue), the AI learned that when it is confused, "blue" is a safe bet. It's like a student who, when they don't know the answer to a history question, just writes "Blue" because they saw that word a lot in their textbook.

The Role of Data Quality
The researchers also looked at how the AI was trained. They compared models trained on massive, messy piles of internet data (like the LAION dataset) with models trained on smaller, carefully cleaned and curated datasets.

They found that the messy, huge datasets didn't actually make the AI smarter at colors. In fact, models trained on the carefully curated, high-quality data were much better at matching human colors. The sheer size of the data didn't matter as much as the quality. It's like studying for a test: reading a million random, unedited blog posts might give you a lot of facts, but reading a well-edited textbook helps you understand the concepts better. The "clean" models made fewer mistakes and were less likely to crash into that default blue color.

The Takeaway
So, what does this all mean? The paper shows that while AI is getting really good at recognizing physical objects and their colors, it still struggles with the messy, cultural, and emotional side of human perception. When the AI gets confused, it doesn't admit uncertainty; it just defaults to a "safe" color, usually blue, based on what it saw most often in its training data.

The researchers suggest that to make AI truly understand us, we can't just feed it more data. We need to teach it to handle uncertainty without crashing, and we need to use better, cleaner data to help it learn the subtle, cultural rules of color that humans follow naturally. The game of "Hues and Cues" proved that even the smartest AI models still have a long way to go before they can truly see the world the way we do.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →