← Latest papers
💬 NLP

Model Confidence Under Answer-Preserving Attacks: An Informativeness-Manipulability Frontier

This paper demonstrates that confidence readouts in deployed vision-language systems are highly vulnerable to answer-preserving attacks, proving that they function as integrity-sensitive rather than robust oversight signals capable of being manipulated to accept incorrect answers or degrade overall accuracy below baseline levels.

Original authors: Reza Khanmohammadi, Ivan Brugere, Simerjot Kaur, Charese H. Smiley, Kundan Thind, Mohammad M. Ghassemi

Published 2026-08-10
📖 4 min read☕ Coffee break read

Original authors: Reza Khanmohammadi, Ivan Brugere, Simerjot Kaur, Charese H. Smiley, Kundan Thind, Mohammad M. Ghassemi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are the captain of a spaceship, but you can't see the stars or the controls. Instead, you have a very smart robot co-pilot who looks at the instruments and tells you, "I'm 90% sure we're on the right path," or "I'm only 50% sure, maybe we should stop." You trust the robot's confidence score to decide whether to keep flying or to pause and ask a human for help. This is how many modern AI systems work today: they don't just give you an answer; they give you a "confidence score" to tell you how much you should trust that answer.

But here is the tricky part: what if someone could trick the robot into changing its mind about how sure it is, without actually changing the answer it gives? It's like a magician making a scale tip from "heavy" to "light" while the weight on the scale stays exactly the same. If the robot's confidence score is the only thing stopping you from making a bad decision, and that score can be faked, then your safety net has a hole in it. This paper dives into exactly that hole, asking a scary but important question: Can we trick these AI robots into lying about their confidence while they keep saying the exact same thing?

The researchers in this study decided to play a high-stakes game of "find the weakness" with several powerful vision-language models (AI that can see pictures and answer questions). Their goal was to see if they could create a tiny, almost invisible change to an image that would make the AI's confidence score swing wildly, even though the AI's final answer remained perfectly identical, down to every single letter and space. They treated the AI's confidence score like a fragile glass window: they wanted to see if they could shatter the glass (the confidence) without breaking the frame (the answer).

What they found is that the glass is much more fragile than anyone hoped. Using a method called "answer-preserving attacks," they managed to manipulate the AI's internal confidence scores in almost every test they ran. In fact, they found that for many of the AI models they tested, they could trick the system into accepting a wrong answer as if it were correct, or reject a correct answer as if it were wrong, simply by tweaking the image just enough to fool the confidence meter.

The most surprising part of their discovery is that this isn't just about the image itself. The researchers showed that they could also "push" the AI's internal brain (its hidden states) directly, like nudging a thought in a specific direction, and get the same result without even touching the image. This suggests that the problem isn't just a glitch in how the AI sees pictures, but a deeper issue with how the AI calculates its own certainty.

They also tested several "defense" strategies—ways to try to make the confidence score more robust, like adding noise to the image or training the AI to be less sensitive to changes. Unfortunately, none of these defenses worked well enough to stop the trick. In their simulations, when they used these manipulated confidence scores to make real-world decisions, the system actually performed worse than if they had just ignored the confidence score entirely and accepted every answer.

So, what does this mean for the future? The paper concludes that we cannot blindly trust the AI's internal confidence score as a safety signal, especially if the person feeding the AI data might be trying to trick it. The confidence score is "integrity-sensitive," meaning it can be easily corrupted. The authors suggest that if we want to use these confidence scores to make important decisions, we need to build new, external ways to verify them, rather than relying on the AI to tell us how sure it is. It's a reminder that in the world of AI, just because a robot says "I'm sure," doesn't mean it actually is.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →