← Latest papers
💻 computer science

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

This paper reveals that Multimodal Large Language Models successfully encode coarse-grained visual evidence but struggle to reliably control their reliance on it versus language priors, a limitation that can be mitigated through supervised fine-tuning and activation steering.

Original authors: Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan, Xi Liu, Serge Belongie, Vésteinn Snæbjarnarson

Published 2026-07-30
📖 4 min read☕ Coffee break read

Original authors: Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan, Xi Liu, Serge Belongie, Vésteinn Snæbjarnarson

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are talking to a super-smart robot that has read almost every book ever written and watched millions of hours of movies. This robot is a "Multimodal Large Language Model" (MLLM). It's like a genius librarian who knows everything about the world from its training data—its "prior knowledge." But this robot also has eyes. It can look at a picture and combine what it sees with what it knows. The big question scientists have been asking is: When the robot looks at a picture that contradicts what it knows (like a picture of a horse with five legs), does it fail because it can't see the extra leg, or does it fail because it sees the leg but decides to ignore it? This paper dives into that mystery, treating the robot's brain like a detective story to figure out where the breakdown happens: in the "senses" (seeing) or in the "decision-making" (using what it sees).

The researchers, led by Jiaang Li and colleagues, set out to solve this puzzle using a clever two-part investigation. First, they wanted to know if the robot was actually "seeing" the weird stuff in the pictures. To test this, they didn't just ask the robot questions; they tried to rebuild the pictures from the robot's internal brain signals. They found that even when the robot gave the wrong answer, the "blueprint" of the weird picture (like a horse with five legs) was still clearly visible in its final thoughts. It was as if the robot's eyes were working perfectly, but its brain was choosing to close them. This ruled out the idea that the robot was just blind to the details.

Next, the team created a new game called "WhatIfVis" to test how well the robot could switch between trusting its eyes and trusting its memory. They showed the robot pictures that broke the rules of nature (like a red watermelon) and asked it to answer questions in two ways: "Look only at the picture" or "Ignore the picture and use what you know." They found that without extra training, the robot was a bit of a mood swing. Sometimes it ignored the picture and gave the textbook answer (like saying a watermelon is green even when it's red). Other times, it got so stuck on the picture that it couldn't follow instructions to ignore it. The robot wasn't broken; it just didn't have a reliable "switch" to decide when to look and when to think.

However, the good news is that the researchers found the switch. By giving the robot a little bit of practice (called Supervised Fine-Tuning) on just one type of tricky picture, they taught it how to control its attention. Even better, they discovered a specific "knob" inside the robot's brain—a single mathematical direction—that could turn the robot's focus on or off without needing any verbal instructions at all. When they turned this knob, the robot became much better at following the rules, whether it was looking at the picture or ignoring it.

The study also compared how the robot handled pictures versus written sentences. They found that it was much easier to make the robot follow a written sentence that broke the rules than a picture. It's like the robot is a much better listener than a viewer. This gap actually got bigger as the robots got smarter and larger, suggesting that as these models grow, they might get even more stubborn about ignoring what they see.

In the end, the paper suggests that for the big, obvious things we study (like color, size, and how many legs an animal has), these AI models aren't failing because they can't see. They fail because they haven't learned a consistent way to decide when to trust their eyes and when to trust their memory. But since the researchers found a specific "knob" that controls this behavior, it means we might be able to fix it in the future, teaching these digital brains to be more flexible and reliable when the world doesn't look exactly like their textbooks.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →