What Color Is the Text? A Benchmark for Hallucination Induced by Image-Embedded Prompt
This paper introduces the "Embedded Stroop" benchmark and the "What-Color-Is-the-Text" (WCIT) dataset to demonstrate that Multimodal Large Language Models frequently suffer from "Stroop hallucinations," where they incorrectly answer with the semantic meaning of text embedded in an image rather than the actual color of that text, revealing that semantic legibility can override visual color perception.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of artificial intelligence, there is a growing class of systems known as multimodal large language models. These are the digital minds that can see images and read text, then combine those two senses to answer questions about the world. They have become remarkably good at describing a photo or solving a puzzle based on a picture. However, just like humans, these systems are not immune to confusion. They can sometimes get tricked by what they see, or more accurately, by how they interpret what they see. A classic example of this confusion comes from a famous psychological test called the Stroop effect. In this test, a person is shown a word like "RED" printed in blue ink. If asked to name the color of the ink, the human brain often stumbles because it automatically reads the word "RED" before it can focus on the blue color. The brain has to work hard to ignore the meaning of the word and focus only on the visual color. For years, scientists have wondered if these advanced computer models suffer from the same kind of mental glitch when they look at images.
A team of researchers at Beihang University in China decided to put this question to the test with a new experiment designed specifically for machines. They created a benchmark called "What Color Is the Text," which presents the computer with a simple but deceptive image. Instead of asking the computer a question in a separate text box, they wrote the question directly onto the image itself. Imagine a white square with the words "What color is the text?" written in black ink. Inside that same image, there is another word, such as "BLUE," but this word is printed in red ink. The computer is asked to identify the color of the ink used for the word "BLUE." The challenge is that the computer must ignore the meaning of the word "BLUE" and report that the ink is actually red. This setup forces the machine to choose between what the text says and what the eyes see.
When the researchers ran this test on sixteen different artificial intelligence models, ranging from open-source projects to powerful commercial systems, the results were striking. The machines failed far more often than they succeeded. In fact, when the question was embedded directly into the picture, the models frequently ignored the actual color of the ink and simply answered with the color named in the text. For instance, if the word "GREEN" was written in purple ink, the model would confidently say "green." This happened even though the models could correctly identify the color if the question was asked in a separate text box rather than written on the image. The study found that this specific type of error, where the machine gets distracted by the meaning of the words rather than the visual reality, occurred in a significant portion of the attempts. In some cases, more than half of the errors made by the models were due to this specific confusion, where they prioritized the text over the color.
The researchers wanted to be sure this wasn't just a case of the models being bad at naming colors in general. To check this, they looked at a group of images where the models had previously succeeded in naming colors correctly when the question was asked separately. Even when the models knew the color names, they still fell into the trap when the question was embedded in the image. This proved that the problem was not a lack of vocabulary or a simple inability to see colors, but a deeper issue where the machine's ability to read text overpowered its ability to see visual details. The study also tested what would happen if they made the text harder to read. When they flipped the image upside down or covered parts of the words with black blocks, the models made fewer mistakes. This suggests that the confusion comes from the machine's ability to clearly read the semantic meaning of the words. When the text is scrambled or hidden, the machine relies more on the visual color, and its performance improves.
Despite these findings, the researchers noted that the problem is not easily solved by simply telling the machine to pay attention. They tried giving the models special instructions to warn them about the trick, and while this reduced the number of times the models fell for the text, it did not make them much better at actually seeing the correct color. The models simply started making different kinds of mistakes. The study concludes that current artificial intelligence systems have a fundamental weakness: when visual information and text conflict, the text often wins. This suggests that these models are not truly "seeing" the world in the way humans do, but are instead heavily influenced by the patterns of language they have learned. As these systems are deployed in real-world situations, such as reading road signs or identifying objects in complex environments, this tendency to trust text over sight could lead to unreliable decisions. The work serves as a clear warning that for these machines to become truly robust, they need to learn how to balance what they read with what they see, rather than letting one sense completely dominate the other.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.