← Latest papers
🤖 machine learning

TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint

This paper introduces TRAPSBench and the PECS metric to demonstrate that while Vision-Language Models possess the internal capability to detect when visual evidence is insufficient for a decision, they consistently fail to express this epistemic restraint, revealing a critical bottleneck in their output generation rather than their perception.

Original authors: Fnu Pramono, John Cai, Sourabh Kulkarni

Published 2026-08-14
📖 3 min read☕ Coffee break read

Original authors: Fnu Pramono, John Cai, Sourabh Kulkarni

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to play a game of pool. You show it a video of the balls rolling, and you ask, "Which pocket will the eight-ball go into?" If the video is clear, the robot should calculate the physics and give you the answer. But what if you cover the table with a giant, opaque tarp halfway through the video, or if the balls are bouncing so wildly that no one could possibly predict where they land? A truly smart robot shouldn't just guess; it should raise its hand and say, "I can't see enough to know." This ability to know what you don't know is called epistemic restraint. It's the difference between a confident liar and a cautious expert. In the world of Artificial Intelligence, specifically Vision-Language Models (VLMs)—computers that can see videos and talk about them—researchers have been wondering: Do these models actually know when they are in the dark, or do they just pretend to see through the walls?

This paper, titled TRAPSBench, dives into that exact question. The researchers built a special video game for AI called TRAPSBench, which features 1,404 pairs of physics scenarios. In each pair, one video is a "control" where the outcome is clear and easy to predict, and the other is a "void" where the answer is impossible to know because of a hidden wall, a chaotic mess, or a question that makes no sense. They then tested 16 different AI models to see if they would admit defeat when the answer was hidden.

Here is the twist: The models are actually very good at knowing the answer is hidden, but they are terrible at saying they don't know. It's like a student taking a test who secretly knows they don't have the right answer, but because they are so eager to please the teacher, they write down a confident, made-up answer anyway. The researchers found that when they probed the AI's internal "brain" (its hidden states), they could detect with high accuracy (up to 91% on a scale called AUROC) whether the model realized the evidence was missing. However, when the model actually spoke its answer, it ignored that internal warning and just guessed.

The study also discovered a funny imbalance: these AI models are much better at spotting when a question is nonsense in text (like asking "What is the weight of the color blue?") than when the visual evidence is missing. They can detect a text-based impossibility about four times faster than a visual one. Furthermore, the researchers tried to "steer" the AI by nudging its internal signals. They found that if they pushed the AI in a specific direction, they could force it to stop guessing and say "I don't know," proving that the ability to restrain itself was already there, just locked behind a door the model refuses to open on its own.

In short, the paper suggests that the problem isn't that these AI models are blind or stupid; it's that they are over-eager. They encode the truth that they are unsure, but their output is blocked by a "gate" that forces them to answer anyway. The researchers conclude that to make these AI systems truly reliable, we probably need to fix the part of the system that decides what to say, rather than trying to teach them how to see better.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →