← Latest papers
💬 NLP

System-Mediated Attention Imbalances Make Vision-Language Models Say Yes

This paper identifies that vision-language model hallucinations, specifically the "yes-bias," are driven by imbalanced attention allocation toward redundant system weights rather than just image or text, and demonstrates that causally redistributing this attention back to the input modalities effectively mitigates the bias.

Original authors: Tsan Tsai Chan, Varsha Suresh, Anisha Saha, Michael Hahn, Vera Demberg

Published 2026-04-27
📖 4 min read☕ Coffee break read

Original authors: Tsan Tsai Chan, Varsha Suresh, Anisha Saha, Michael Hahn, Vera Demberg

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery by looking at a photo and reading a witness statement. To solve it, you need to balance three things: your eyes (the image), your brain’s logic (the text), and your internal habits (the "system").

This paper explores why AI models (specifically Vision-Language Models) often fail at this. They suffer from something called "Yes-Bias"—a glitch where, no matter what the photo shows, the AI just reflexively shouts, "Yes!"

Here is the breakdown of how they discovered why this happens and how to fix it.

1. The Problem: The "Lazy Detective" Syndrome

Most researchers thought AI hallucinated because it wasn't "looking" hard enough at the picture. They called this the Image-Centric Hypothesis. They thought the AI was like a detective who keeps their eyes closed while reading the report.

However, these researchers proposed a different idea: the System-Mediated Hypothesis. They argued the problem isn't just that the AI isn't looking at the picture; it’s that the AI is too distracted by its own "internal autopilot."

The Analogy: Imagine a detective who has a massive, heavy manual of "Standard Operating Procedures" (the System) sitting on their desk. Every time they try to look at a clue (the Image) or read a note (the Text), they get distracted by flipping through the pages of that manual. They spend 70% of their energy on the manual and only 30% on the actual crime scene. Because they are so distracted by their "habits," they stop being precise and just start guessing "Yes" to everything to get the job done faster.

2. The Discovery: The "Attention Sink"

The researchers found that in the final stages of "thinking," the AI spends a massive amount of its "attention" on meaningless tokens (like the start of a sentence or formatting marks). They call these Attention Sinks.

These sinks act like mental black holes. They suck up all the AI's focus, leaving very little "brainpower" left to actually process the fine details of the image or the specific words in the question. When the AI runs out of focus, it stops being a scientist and starts being a "yes-man."

3. The Fix: "Redistributing the Focus"

To prove this, the researchers performed a "brain surgery" on the AI. They didn't just tell the AI to "look harder at the image." Instead, they took the attention being wasted on the "System" (the manual) and forcefully redistributed it back to the Image and the Text.

The Analogy: It’s like taking the heavy manual away from the detective and saying, "Stop reading the rules for a second and actually look at the fingerprints and the witness's words!"

The Result? It worked! By taking focus away from the "autopilot" and giving it back to the actual evidence, the AI stopped blindly saying "Yes" and started giving much more accurate answers.

4. Why does this matter?

The researchers found that this "Yes-Bias" is most dangerous in compositional tasks—tasks that require seeing tiny differences (e.g., "Is there a strawberry on this pizza?" vs. "Is there a pepperoni on this pizza?").

When the AI is distracted by its "System" habits, it only sees a "blurry, general idea" of a pizza. It thinks, "I see a pizza... pizzas are food... I'll just say Yes!" By fixing the attention imbalance, the AI is forced to look at the "fine-grained" details, allowing it to see the difference between a strawberry and a pepperoni.

Summary in a Nutshell

  • The Old View: The AI is blind (needs more image attention).
  • The New View: The AI is distracted by its own internal habits (needs less system attention).
  • The Solution: Stop letting the AI obsess over its own "autopilot" and force it to spend that energy on the actual visual and textual clues.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →