← Latest papers
🤖 AI

When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence

This paper demonstrates that open vision-language models fail to make optimal decisions about when to ask humans for help due to their reliance on prompt formatting rather than actual sensor evidence, but their diagnostic accuracy and decision-making can be significantly improved by incorporating diverse sensor data and grounding their actions in measured reliability and task costs.

Original authors: Eshika Pathak, Leela Krishna

Published 2026-09-21
📖 5 min read🧠 Deep dive

Original authors: Eshika Pathak, Leela Krishna

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

When a robot working alongside a person drops a cup or fails to pick up a tool, a critical moment arrives before anyone speaks. The machine must decide its next move: should it try to fix the problem on its own, check its own internal sensors for clues, or stop and ask the human for help? This decision is not just about politeness; it is a calculation of risk. Asking for help costs time and interrupts the human's focus, but acting blindly on a wrong guess can lead to more damage. For years, researchers have assumed that if a robot can see a failure, it can understand why it happened, or that it should simply ask for help whenever it feels unsure. However, a new study challenges these assumptions, suggesting that the robot's ability to diagnose a problem depends entirely on which sensors it uses, and that current artificial intelligence models are surprisingly bad at knowing when they should stop and ask.

To investigate this, researchers built a controlled environment where they could create specific failures with known causes. They programmed a simulated robot arm to perform a simple task: pick up an object and place it somewhere. They then introduced three distinct types of problems. First, the robot might fail to see the object because it was hidden, missing, or the lighting was too dim. Second, the robot might try to lift the object but drop it because its grip was too weak, the object was too heavy, or the surface was too slippery. Third, the robot might fail to place the object because the target container was full or moved out of reach. Because the researchers created these failures themselves, they knew the exact cause of every mistake, allowing them to test whether a robot could actually figure it out.

The researchers first tested what information different sensors could provide. They found a sharp divide in what the robot could learn. When the failure was about not seeing the object, a camera was excellent at diagnosing the problem, correctly identifying the cause nearly every time. However, when the failure was about dropping the object, the camera was almost useless. No matter how advanced the image analysis was, the camera could not tell the difference between a weak grip, a heavy object, or a slippery surface, because these physical forces are invisible to a lens. In contrast, when the researchers gave the robot its own force data—measurements of how hard the gripper squeezed and how much weight it felt—the robot could diagnose the drop with near-perfect accuracy. This revealed a fundamental truth: a robot cannot diagnose a problem if the information about that problem is not present in the data it is looking at.

The study then turned to six different open-source artificial intelligence models, the kind of systems often used to give robots the ability to understand language and images. The researchers asked these models to look at the camera footage of a failed attempt and guess what went wrong. The results were startling. The models did not behave like cautious scientists who know when they are missing information. Instead, their behavior was dictated by the layout of the question they were asked. If the option to say "I don't know" was listed last, the models would almost always refuse to answer, claiming they could not determine the cause. If the researchers moved that same option to the top of the list, the models suddenly stopped refusing and started guessing, often with high confidence, even when they were wrong. Their confidence levels did not reflect reality; a model could be 100 percent sure of a wrong answer, or 50 percent sure of a right one, with no pattern to connect the two.

The researchers also tested whether giving the models the force data, written out as simple text, would help. For four of the six models, this extra information made a massive difference. Suddenly, they could diagnose the drop failures correctly, moving from guessing randomly to getting the answer right more than half the time. This proved that the models were not inherently incapable of understanding the physics of the failure; they were simply missing the right sensor data. However, even with this new ability, the models still failed to make the right decision about when to ask for help. They did not ask more often when the question was cheap and the risk of acting alone was high, nor did they stop asking when the question was expensive. Their strategy for asking was rigid and ignored the cost of interrupting a human.

The most effective strategy for a robot, according to the study, is not to rely on the model's own feeling of confidence, which is often misleading. Instead, the decision to ask should be based on two measurable facts: how accurate the robot's diagnosis actually is, and how costly it is to ask a human. If the robot's sensors can reliably tell it what is wrong, it should act. If the sensors cannot provide the answer, or if the robot's own diagnosis is likely to be wrong, it should ask. The study shows that a single question to a human can lift a robot's success rate from near zero to over 80 percent, but only if the robot asks at the right moment. Today, the choice to ask is often a guess, but the path forward is clear: robots need to be wired to know the limits of their own senses and the true cost of a question, rather than trusting the confidence of a machine that does not know what it does not know.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →