When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models
This paper reveals that task-irrelevant text in multimodal large language models systematically biases binary visual judgments by inducing a predictable affine transformation in decision margins, characterizing the effect as an estimable geometric distortion rather than unstructured noise.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern landscape of artificial intelligence, a new generation of systems has emerged that can see and speak at the same time. These multimodal models are trained to look at an image and answer questions about it, or to describe a scene in detail. They are powerful tools, capable of identifying objects, reading text within pictures, and reasoning about complex situations. However, these systems do not operate in a vacuum. In real-world applications, they are often fed extra information alongside the image: a user's previous messages, a retrieved article, or a database entry. This extra text is meant to help the model understand the picture better. But what happens when that extra text has nothing to do with the image? Does the model ignore the noise, or does it get confused? This question sits at the heart of a recent investigation into how these intelligent systems process information when their attention is divided between a visual scene and unrelated words.
Researchers set out to test exactly this scenario using a controlled experiment. They took a series of images and questions, such as asking whether a dog in a photo was jumping over a hurdle. They then created two versions of the input for each question. In the first version, the model saw only the image and the question. In the second version, the model saw the same image and question, but with a block of completely irrelevant text inserted into the prompt. This text was a random sentence from a book or article, something like a story about a person traveling to France, which had no connection to the dog or the hurdle. The researchers wanted to see if this random noise would change the model's mind.
The results were striking. When the irrelevant text was added, the models did not simply ignore it. Instead, they consistently changed their answers. More importantly, they did not change their minds in a random way. The models became significantly more likely to say "No" to questions, even when the image clearly showed a "Yes" situation. For example, if a model was confident that a dog was jumping, the addition of unrelated text often made it hesitate and flip its answer to say the dog was not jumping. This shift happened across different types of images and different models, suggesting a systematic flaw rather than a random glitch. The models were not just getting slightly less accurate; they were being pushed toward a specific type of error, one where they doubted the visual evidence they were looking at.
To understand why this was happening, the researchers looked deeper than just the final answers. They examined the internal confidence scores the models generated before making a decision. They found a very specific pattern in how the models processed the information. When the models saw the image alone, they had a certain level of confidence in their answer. When the unrelated text was added, that confidence did not disappear or become chaotic. Instead, it shifted in a predictable, mathematical way. The new confidence level was a distorted version of the original, compressed and pushed in a specific direction. It was as if the extra text acted like a filter that squeezed the model's certainty and tilted the scale against the visual evidence.
This discovery ruled out the idea that the models were simply confused by the extra words or that the text acted as random static noise. If it were random noise, the errors would have been scattered and unpredictable. Instead, the errors followed a strict rule. The researchers described this rule as a consistent shift, where the model's internal judgment was stretched and moved by the presence of the text. They found that this shift could be measured and even predicted. By understanding the pattern of the shift, they were able to create a simple correction method. This method could partially undo the damage caused by the irrelevant text, restoring the model's ability to trust what it saw in the image.
The study also revealed that the way the text was presented mattered. When the irrelevant text was framed as if it might be related to the image, the distortion was even stronger. The models seemed to treat the random words as if they were clues, trying to fit them into their understanding of the picture, which led to even greater confusion. This suggests that the models are not just passive receivers of information but are actively trying to make sense of everything they are given, even when that information is useless. They struggle to distinguish between what is relevant to the visual task and what is just background noise.
These findings offer a new way to look at how artificial intelligence handles information. It turns out that adding extra text to a visual task does not just add volume; it changes the shape of the model's thinking. The models are sensitive to the context they are given, and that sensitivity can lead them to doubt their own eyes. While the researchers have shown that this effect can be measured and partially corrected, the underlying issue remains: these systems are not yet robust enough to ignore the irrelevant. As these technologies are integrated into more complex real-world systems where they will constantly receive streams of data, understanding how they react to noise is crucial. The work provides a clear diagnostic tool for spotting these biases and a foundation for building systems that can focus on the image, even when the world around them is full of distractions.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.