Are Reasoning Vision-Language Models Robust to Semantic Visual Distractions?
This paper introduces Distract-Bench, a new benchmark demonstrating that Reasoning Vision-Language Models are significantly more vulnerable to semantic visual distractions—meaningful but irrelevant cues—than to traditional perceptual corruptions, as these distractions often mislead the models' reasoning processes despite correct visual perception.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: Smart Students vs. Distracting Classrooms
Imagine you have a very smart student (a Reasoning Vision-Language Model or VLM). This student is great at looking at a picture and a question, thinking through the problem step-by-step, and giving the right answer. They are like a math whiz who can solve complex geometry problems just by looking at a diagram.
For a long time, researchers tested how "tough" these students were by making the classroom messy. They would:
- Blur the blackboard (blur).
- Turn off the lights (noise).
- Shake the camera (weather effects).
This is called Perceptual Robustness. It asks: "Can the student still see the answer if the view is fuzzy?"
This paper introduces a new, sneakier problem.
Instead of making the view fuzzy, the researchers kept the blackboard crystal clear. But, they added a distractor: a piece of paper taped to the board with a true fact that has nothing to do with the question.
- The Question: "How many apples are in the basket?"
- The Image: A basket with 3 apples.
- The Distractor: A bright red sign taped next to the basket that says, "There are 5 oranges in the next room." (This is a true fact, but it's irrelevant).
The researchers call this Semantic Distraction. They wanted to see: If the student can see clearly, will they get tricked by the irrelevant information?
The Experiment: Distract-Bench
The team built a test called Distract-Bench.
- The Setup: They took 506 real-world questions and images (like math problems, charts, and daily scenes).
- The Trick: They used AI and human experts to add "distractors" to the images. These distractors were:
- Factually True: The sign really did say "5 oranges."
- Visually Plausible: It looked like it belonged on the board.
- Totally Irrelevant: It didn't help answer the question about the apples.
- Answer-Preserving: The correct answer (3 apples) didn't change.
They then tested 10 different models (some are "reasoning" models that think step-by-step, and some are standard models) to see how they handled these tricks.
The Surprising Findings
The results were counter-intuitive and revealed a hidden weakness in "smart" models.
1. The "Blur" Test vs. The "Distractor" Test
- When the image was blurry (Perceptual Robustness): The "smart" reasoning models behaved almost exactly like their "dumber" base versions. If the base model couldn't see through the blur, the smart model couldn't either. They inherited the same visual limitations.
- When the image had a distractor (Semantic Robustness): The "smart" reasoning models actually performed worse than the base models! Even though they were better at thinking, they were more easily tricked by the irrelevant facts.
Analogy: Imagine a student who is great at math but gets distracted by a fly buzzing around the room. A simpler student might just ignore the fly and focus on the numbers. The "smart" student, however, starts thinking, "Wait, is that fly a symbol? Does it mean I should add 1 to the answer?" Their ability to over-think actually hurt them.
2. The Reasoning Trap
The researchers looked at why the models failed. They found that the "smart" models didn't just guess wrong; they incorporated the distractor into their logic.
- The Failure Mode: The model would see the red sign saying "5 oranges," treat it as evidence, and build a long, logical chain of reasoning based on that wrong clue.
- The Result: They would produce a very confident, well-explained, but completely wrong answer. The "reasoning" part of the model amplified the mistake.
3. The "Harmful Reference"
The study measured how often the models mentioned the distractor in their final answer.
- Smart models mentioned the irrelevant facts much more often than the base models.
- When they got the answer wrong, they were almost always referencing the distractor as proof for their mistake.
The Conclusion
The paper concludes that being good at reasoning doesn't make you immune to distractions. In fact, for these AI models, the ability to generate long chains of thought can sometimes make them more vulnerable to irrelevant information.
- Old View: We only need to worry if the image is blurry or noisy.
- New View: We also need to worry if the image contains "too much truth." If a model sees a fact that is true but irrelevant, it might grab onto it and let it derail its entire thought process.
The Takeaway: To make these AI models reliable in the real world, we can't just teach them to see better; we have to teach them to ignore things that don't matter, even if those things look important and are factually true.
(Note: This explanation is based strictly on the findings presented in the paper. The authors do not discuss specific clinical applications or future commercial uses in this text, so none are included here.)
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.