Which Source Wins? Task-Dependent Reliance in Vision-Language Models
This paper reveals that Vision-Language Models dynamically shift their reliance between text and images based on task type and evidence structure, showing a preference for text over degraded images in arithmetic problems but a stronger preference for images over degraded text in chart-based reasoning tasks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Human beings are remarkably good at making sense of a confusing world. When we receive information from different senses, such as sight and sound, our brains automatically weigh them. If one source becomes noisy or hard to understand, we instinctively shift our attention to the clearer one. This ability to recalibrate our trust is a fundamental part of how we navigate reality. Modern artificial intelligence systems known as vision-language models are designed to do something similar. These programs process both images and text together, attempting to solve problems or answer questions by combining what they see with what they read. For a long time, researchers assumed these systems simply had a fixed preference, perhaps trusting text more than pictures, or vice versa. But it remained unclear whether these models could dynamically adjust their trust when the two sources disagreed and one of them became difficult to read.
A team of researchers set out to test this flexibility. They wanted to know if a model would change its mind when the text became blurry or the image became garbled, and if it would react the same way regardless of which source was damaged. To find the answer, they created a series of controlled puzzles. They took math problems and rendered them as images, then paired each image with the text of a completely different math problem. This forced the model to choose between two conflicting answers: one derived from the picture and one from the words. They repeated this process with charts and written reports, creating a new dataset where a graph showed one answer and the accompanying text suggested another. In every case, they systematically degraded one source, making it harder to read by adding blur, noise, or missing letters, while keeping the other source perfectly clear. They then watched to see which source the model followed as the confusion grew.
The results revealed a surprising twist. When the researchers tested the models on the math problems, where the conflict was essentially between two versions of text (one printed on a screen and one written out), the models behaved in one specific way. As the text became harder to read, the models shifted their reliance away from the text more strongly than they did when the image was degraded. However, when they tested the same models on the charts and reports, the behavior flipped completely. In this scenario, as the chart became harder to read, the models shifted their reliance away from the visual source more strongly than they did when the text was degraded. This reversal was consistent across six different open-source models and was also observed in two of the most advanced commercial models available. The researchers found that the models did not have a permanent bias toward either seeing or reading. Instead, their reliance on a specific source depended entirely on the type of task and the structure of the evidence presented.
This discovery challenges the idea that these artificial intelligence systems have a single, unchanging way of processing information. The study showed that the models are sensitive to the context of the problem. When the visual information was a simple rendering of text, the models relied more heavily on the text source, shifting away from it only when it became significantly harder to read. But when the visual information was a complex chart, the models shifted away from the visual source more aggressively than they did from the text source. The researchers confirmed that this pattern held true even when they replaced the charts with plain tables of numbers, proving that the effect was not just about the graphical style of the chart but about how the model perceived the relationship between the visual and textual evidence.
The study also addressed whether the models were simply reacting to the amount of information lost. The researchers measured how much the models' accuracy dropped when they were given only the degraded source, and they adjusted their analysis to account for this loss. Even after this careful calibration, the reversal in behavior remained. This suggests that the models are not just reacting to the severity of the damage but are making a more complex judgment about which source is more likely to be correct in a given situation. The findings indicate that the way these models arbitrate between conflicting sources is a setting-dependent behavior, shaped by the task at hand, the nature of the evidence, and the specific model architecture.
By providing a new way to measure how these systems shift their attention, the researchers have offered a clearer picture of how artificial intelligence handles conflicting information. The work demonstrates that these models are not rigid in their preferences but are capable of dynamic reallocation. However, this flexibility is not uniform; it changes based on the context. The study concludes that understanding how these models decide what to trust requires looking at the specific details of the task and the evidence, rather than assuming a universal rule. This insight is crucial for anyone hoping to deploy these systems in real-world scenarios where information might be imperfect or contradictory, as it highlights that the model's trust is a fluid variable, not a fixed setting.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.