← Latest papers
💬 NLP

Small Vision-Language Models Know When They Are Wrong But Cannot Say So: A Two-Model Study of Stated versus Internal Confidence Under Realistic Image Degradation

This study reveals that while small vision-language models possess accurate internal self-knowledge about their errors under realistic image degradation, they fail to verbalize this uncertainty, making internal token probability a superior signal for deferral than stated confidence, though both signals become unreliable under severe low-light conditions.

Original authors: M M Asif Ferdous

Published 2026-07-27
📖 5 min read🧠 Deep dive

Original authors: M M Asif Ferdous

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking through a bustling city with a very smart, but tiny, robot guide. This robot has a camera for eyes and a voice for speaking. Its job is to look at things around you and tell you what they are. But here's the catch: the city is messy. Sometimes the sun is too bright, sometimes it's pitch black, and sometimes your camera lens is smudged or shaking. In the world of artificial intelligence, these "messy" pictures are called degraded images.

When a robot like this gets confused by a bad picture, it needs to know when to say, "I don't know, ask a human," instead of confidently guessing the wrong answer. This ability to know when it is unsure is called uncertainty. There are two ways a robot can show this uncertainty. The first is verbalized confidence: the robot simply says out loud, "I am 90% sure!" The second is internal confidence: a secret, silent signal inside the robot's brain that measures how shaky its thoughts are, even if it doesn't say a word. Most people assume that if a robot says it's confident, it probably is. But what if the robot is just a bad liar? That is the big question this paper asks.


The Story of the Silent Signal

In this study, researchers put two small, open-source robot guides (called Qwen2-VL and SmolVLM) through a rigorous test. They didn't just show them clean, perfect photos. Instead, they took 100 pictures of food and deliberately ruined them in six different ways: squashing the quality like a JPEG file, blurring them with motion, dimming the lights until they were almost dark, adding blinding glare, tilting the camera, or shrinking and stretching the image. They did this at three levels of severity for each type of ruin.

The robots had to guess what the food was in every single ruined picture. Then, the researchers checked two things:

  1. Did the robot say, "I'm confident!" or "I'm not sure"? (The Verbalized Signal)
  2. What was the robot's secret, internal score of how likely it was to be right? (The Internal Signal)

The Big Surprise: The Robot That Won't Admit It's Wrong

The results were a bit like watching a magician who is terrible at their job but keeps bowing and saying, "That was perfect!"

The Verbalized Signal (The Robot's Mouth):
When asked to state their confidence, the robots were incredibly stubborn. No matter how terrible the picture was, no matter how blurry or dark it got, they kept saying they were about 87% to 90% sure. Even when the picture was so dark the robot was basically guessing randomly (getting the answer right only 22% of the time), it still confidently said, "I'm 87% sure!" It was as if the robot had a broken confidence dial stuck on "High." In fact, for one of the robots, asking it to give a confidence number was so difficult that it mostly just refused to answer, giving only the name of the food without any number at all.

The Internal Signal (The Robot's Brain):
However, when the researchers peeked inside the robot's brain to see its secret "token probabilities" (a fancy way of measuring how sure the robot's own math was), they found something totally different. The robot's internal signal was actually very honest! When the picture was clear, the internal score was high. When the picture was blurry or dark, the internal score dropped significantly. This secret signal was so good at spotting mistakes that it could tell the difference between a right answer and a wrong one with 92% to 99% accuracy.

The "Low Light" Trap

There was one scary moment in the story, though. When the researchers made the pictures extremely dark (the "severe underexposure" level), both signals failed. The robots' accuracy crashed (dropping from 99% down to 22%), and even the honest internal signal stopped working. It couldn't tell the difference between a right guess and a wrong one anymore. It was as if the darkness was so thick that even the robot's secret brain fogged up.

What This Means for You

So, what did the researchers learn?

  1. Don't trust the robot's mouth. If you ask a small robot guide how sure it is, it will likely lie and say it's very sure, even when it's totally lost. The paper shows that this "stated confidence" is useless for knowing when to stop and ask a human for help.
  2. Trust the robot's brain (mostly). The secret internal math is much better. It knows when it's confused and can signal that to a computer system, allowing the system to say, "Okay, I'm not sure, let's get a human to check this."
  3. But watch out for the dark. Even the honest internal signal breaks down when the images are too dark. If you are using these robots in a dark room or at night, you can't rely on the robot to tell you it's failing. You need a separate "safety check" to make sure the picture isn't too dark before the robot even tries to look at it.

In short, small AI robots know when they are wrong deep inside their code, but they just won't say it out loud. They are like a student who gets a test question wrong but still raises their hand and says, "I'm totally sure of this!" The only way to catch them is to look at their secret scratch paper, not listen to what they say.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →