← Latest papers
💬 NLP

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

This paper introduces SIGNPOST-Bench, a controlled counterfactual benchmark comprising 25,555 image variants to evaluate how multimodal large language models resolve conflicts between visual and textual cues, revealing that adversarial text interventions significantly degrade geolocation accuracy and expose distinct arbitration behaviors across 20 evaluated models.

Original authors: Sirun Li, Minghao Liu, Ling Dai, Yong Li, Haoxin Lyu, Junting Zhou, Fan Zhang

Published 2026-08-06
📖 4 min read☕ Coffee break read

Original authors: Sirun Li, Minghao Liu, Ling Dai, Yong Li, Haoxin Lyu, Junting Zhou, Fan Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery: "Where was this photo taken?" You have two clues to help you. The first is the visual scene: the architecture, the trees, the cars, and the way the light hits the buildings. The second is the text written on signs, storefronts, or billboards in the picture. Usually, these clues agree. A sign saying "Paris" sits on a building that looks like it's in Paris. But what happens when they disagree? What if the photo shows a snowy, wooden cabin in the mountains, but a bright neon sign in the window screams "Welcome to the Sahara Desert"?

This is the world of Multimodal Large Language Models (MLLMs). These are super-smart AI computers that can "see" images and "read" text at the same time. Scientists have been teaching them to combine these clues to make guesses about the world. But there's a big question: When the visual clues and the text clues fight each other, which one does the AI trust? Does it ignore the confusing sign and look at the snow? Or does it blindly believe the sign and forget the snow? Until now, we didn't have a good way to test this specific kind of confusion.

Enter SIGNPOST-Bench, a new experiment designed by researchers to play a game of "trick the AI." They took thousands of real-world photos and created a special set of "counterfactual" versions for each one. Think of it like a photo editing workshop where they take one original picture and make four new copies:

  1. The Blank: They erased the text completely.
  2. The Similar: They swapped the sign for a different one that still makes sense for the location (e.g., changing "Welcome to Michigan" to "Welcome to Detroit").
  3. The Random: They put up a sign that has nothing to do with the place (e.g., "Welcome to the Moon").
  4. The Adversarial: This is the tricky one. They replaced the sign with a lie that points to a specific, different location (e.g., changing "Welcome to Michigan" to "Welcome to New York").

The researchers then asked 20 different AI models from seven major tech companies to guess the location for all five versions of every photo. They wanted to see if the AI could spot the lie or if it would get tricked by the new sign.

The results were quite revealing. When the AI looked at the original photos, it was pretty good at guessing the location. But when they introduced the Adversarial lies, the models got confused. On average, the AI's guess became 4.8 times less accurate. Instead of being off by about 282 kilometers, the AI was now off by 1,347 kilometers.

Even more interesting, the AI didn't just get worse; it got misled. When the sign said "New York," the AI didn't just guess randomly; it actually moved its guess closer to New York. In fact, for every single model tested, the presence of the conflicting text pulled the answer toward the lie. About 6.5% to 20.1% of the time, the AI's guess ended up less than 50 kilometers away from the fake location the sign was pointing to.

The study also found that being "smart" at reading signs or seeing pictures didn't guarantee being good at spotting a conflict. Some models that were excellent at guessing locations with clean, honest photos were actually worse at ignoring the lies than some smaller models. It turns out that being able to read a sign doesn't mean you know when to ignore it if it contradicts the rest of the picture.

Finally, the researchers tried to "coach" the AI, asking it specifically to look out for conflicts. While this helped the AI realize something was wrong about 49% of the time, it didn't actually help it guess the correct location any better. The AI knew it was confused, but it still couldn't figure out how to solve the puzzle.

In short, this paper shows that current AI models are still very easily tricked when text and images disagree. They tend to trust the text too much, even when it's a blatant lie. The researchers built this new "SIGNPOST-Bench" test to help future AI developers understand exactly where their models are failing, so they can teach them to be better detectives who know when to trust their eyes and when to question the signs.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →