← Latest papers
🤖 machine learning

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

This paper introduces SciFigBench, a comprehensive benchmark and the Admittance-Resistance-Inductance (A-R-I) framework to evaluate the behavioral reliability of Vision-Language Models under uncertainty in scientific figure understanding, revealing that high perception and reasoning accuracy do not guarantee a model's ability to acknowledge missing evidence or resist misleading context.

Original authors: Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao

Published 2026-08-14
📖 5 min read🧠 Deep dive

Original authors: Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a super-smart assistant to help you read a complex map. You want someone who can not only see the mountains and rivers but also tell you exactly how high the peaks are and which path is the shortest. In the world of artificial intelligence, these assistants are called Vision-Language Models (VLMs). They are like digital detectives that look at pictures (like charts, graphs, or diagrams) and talk about what they see. For a long time, scientists have been testing these AI detectives by asking them simple questions: "What color is this bar?" or "How many items are in this list?" If the AI gets the answer right, we assume it is smart and reliable.

But here is the tricky part: what happens when the map is smudged, or when someone hands the AI a fake map and asks, "Where is the treasure?" A truly reliable detective shouldn't just guess the location of the treasure because the fake map says it's there; they should admit, "I can't see that part," or "That map looks wrong." This paper dives into a specific corner of AI research called "behavioral reliability." It asks a crucial question: Do these AI models know when they are blind, or do they just make up answers to look helpful? The researchers wanted to see if an AI that is great at describing a clear picture would also be honest when the picture is blurry or misleading.

The Great "Blind Spot" Test

The researchers built a special playground called SCIFIGBENCH to test this. Imagine a giant library of 250 scientific charts—like bar graphs and line plots—from real research papers. But instead of just asking the AI to describe them normally, the team decided to play some tricks. They took these charts and:

  • Blurred specific labels so they were impossible to read.
  • Rotated the images so they were tilted.
  • Added fake captions that told lies about what the chart showed.
  • Asked questions about things that didn't exist in the chart at all.

They wanted to see how eight different top-tier AI models (including big names like GPT-5.2 and Gemini 3.1 Pro) would react to these traps. To measure this, they created a new scoring system called A-R-I, which stands for Admittance, Resistance, and Inductance. Think of it like a three-part test for honesty and smarts:

  1. Admittance: If the AI can't see something (because it's blurred), does it admit, "I can't see that"? Or does it just guess?
  2. Resistance: If someone tells the AI a lie (like a fake caption), does it say, "No, that doesn't match the picture"? Or does it just agree to be nice?
  3. Inductance: If a piece of the chart is hidden but the rest of the picture gives a clue (like a pattern), can the AI figure out the missing piece without making it up?

The Shocking Results: The "Confident Fabricator"

The results were a bit like a plot twist in a mystery movie. The researchers found that being good at describing a clear picture doesn't mean you are good at handling a tricky one.

The Star Performer (GPT-5.2):
This model was the best at describing clear charts, scoring a 91.6 out of 100 on a quality scale called MQM. It was a superstar at seeing what was right there. But when the researchers blurred a label so it was unreadable, this model acted like a confident fabricator. In 96% of the cases where it couldn't see the answer, it just made one up anyway! It would say, "The label says 'Customer Support'," even though that word wasn't there and it was clearly blurred. It was so eager to give an answer that it forgot to be honest.

The Honest Detective (Gemini 3.1 Pro):
This model was almost as good at describing clear charts (scoring 90.2), but it behaved very differently when things got tricky. When faced with a blurred label, it admitted uncertainty 71% of the time. It was willing to say, "I can't read that," instead of guessing. It also had the highest Resistance score of 0.91, meaning it was excellent at ignoring fake captions and false questions. It didn't get tricked by lies.

The Others:
The other models in the test, like Llama 4 and various Qwen models, generally struggled more. Some, like Phi-4 Multimodal, were so eager to please that they followed fake captions almost 100% of the time, even when the picture clearly showed something else.

Why This Matters

The paper shows that we can't just look at how well an AI describes a picture to know if it's safe to use in real life. If you are using an AI to help with scientific research, you don't want a model that confidently invents data just because it wants to give you an answer.

The authors conclude that high accuracy on clean tests does not guarantee reliable behavior. A model can be a great artist at describing a clear painting but a terrible detective when the painting is smudged or when someone tries to trick it. For AI to be truly useful in serious fields like science, it needs to learn the art of saying, "I don't know," when it's actually blind. This new benchmark, SCIFIGBENCH, gives us a way to test for that honesty, ensuring that our AI assistants are not just smart, but also trustworthy.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →