← Latest papers
💬 NLP

Real Images, Worse Judgments: Evaluating Vision-Language Models on Concreteness and Imagery

This paper demonstrates that providing real-image contexts to vision-language models often degrades their performance on lexical judgments of concreteness and imagery by introducing sensitivity to spurious visual cues, suggesting a need for better calibration of when visual input should inform such tasks.

Original authors: Yifan Jiang, Ruoxi Ning, Sheng Yao, Freda Shi

Published 2026-05-27
📖 5 min read🧠 Deep dive

Original authors: Yifan Jiang, Ruoxi Ning, Sheng Yao, Freda Shi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are taking a test where you have to guess how "tangible" or "picture-able" a specific word is. For example, is the word "apple" easy to picture? Yes. Is the word "freedom" easy to picture? Not really.

Now, imagine you are taking this test while someone is showing you a random picture next to the word.

This paper investigates what happens when modern AI models (called Vision-Language Models) take this test. The researchers wanted to see if showing the AI a picture helps it understand the word better, or if the picture actually distracts it and makes it give a worse answer.

Here is the breakdown of their findings using simple analogies:

1. The Setup: The "Distracting Classroom"

The researchers gave the AI a list of words and asked it to rate them on a scale of "how concrete" (real/touchable) or "how easy to imagine" they are.

  • The Control Group: The AI saw the word alone (or with a blank white square).
  • The Test Group: The AI saw the word plus a real photo retrieved from the internet.

The Big Surprise:
You might think, "If I show an AI a picture of a 'dog', it should be better at knowing 'dog' is a concrete word." And for simple, concrete words, the AI did okay.

However, for abstract words (like "justice," "freedom," or "nature"), the pictures made the AI worse.

  • The Analogy: Imagine a student taking a math test. If you hand them a picture of a calculator next to the question "What is 2+2?", they might get it right. But if you ask them, "Is the concept of 'honesty' hard or easy to visualize?" and you hand them a picture of a sunset, the student gets confused. They look at the beautiful sunset and say, "Oh, that's very concrete and real!" and give the word "honesty" a high score for being "real," even though the word itself isn't. The picture acted like a loud, distracting neighbor shouting the wrong answer.

2. The "Human vs. Robot" Comparison

To make sure this wasn't just a weird AI problem, the researchers asked actual humans to take the same test.

  • The Result: Humans also got slightly distracted by the pictures, but they didn't fall apart. They knew the picture was just a decoration.
  • The AI Problem: The AI, however, treated the picture as hard evidence. It couldn't tell the difference between "a picture of a dog" (which proves the word 'dog' is real) and "a picture of a sunset" (which has nothing to do with the word 'freedom'). The AI started guessing based on the picture's content rather than the word's meaning.

3. Why Did This Happen? (The "Internal Glitch")

The researchers looked inside the AI's "brain" (its internal data layers) to see what was going wrong.

  • The Shift: When the AI saw a real picture, its internal understanding of the word actually shifted. It was like the AI's brain suddenly decided, "Oh, I'm looking at a picture now, so I should answer based on what I see, not what I read."
  • The Sensitivity: The AI became overly sensitive to "spurious cues." This is a fancy way of saying it latched onto random details in the photo that had nothing to do with the word. For abstract words, this caused the AI to overestimate how "real" the word was.

4. The Fix: The "Teacher's Note"

Since the AI was getting distracted, the researchers tried a simple trick. They added a tiny instruction to the prompt, like a note from a teacher:

"Hey, this picture might not be related to the word. Please ignore the picture and just rate the word itself."

The Outcome:

  • This simple note acted like a mental filter. It told the AI to stop looking at the distracting neighbor and focus on the test question.
  • The AI's performance improved significantly, especially for those tricky abstract words. It didn't fix everything perfectly, but it stopped the AI from making the biggest mistakes.

Summary

The paper concludes that while we often assume giving AI pictures makes it smarter, sometimes the pictures are actually a trap.

  • For concrete things (like a picture of a cat helping you identify the word "cat"), images help.
  • For abstract things (like a picture of a storm helping you identify the word "peace"), images confuse the AI, making it rely on the wrong clues.

The solution isn't to stop using pictures, but to teach the AI when to ignore them. Just like a good student knows when to look at a diagram and when to close their eyes and think about the concept, these AI models need to learn when the visual context is a helpful hint and when it's just noise.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →