← Latest papers
🤖 AI

Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs

This study demonstrates that while evidence-sufficiency prompting significantly reduces unsafe overconfidence in clinical LLMs, the magnitude of this safety gain is highly judge-dependent and often incurs substantial, model-specific costs to diagnostic helpfulness, indicating that such safety improvements should be reported as relative directional trends rather than calibrated absolute rates.

Original authors: Koyar Afrasyab

Published 2026-07-21
📖 4 min read☕ Coffee break read

Original authors: Koyar Afrasyab

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a super-smart robot to be a doctor. You want it to be helpful, giving great advice when it knows the facts. But you also want it to be safe, meaning it shouldn't guess wildly or give a confident diagnosis when it's missing key information, like a fever or a blood test result. In the world of Artificial Intelligence, this is a tricky balancing act. We have these "Large Language Models" (LLMs) that can chat like humans, and we need to test if they are safe enough to help real doctors. To do this, scientists often use a "judge"—another AI that reads the robot's answers and gives it a score for safety. But here's the big question: Is the robot actually changing its behavior to be safer, or is it just learning how to trick the judge into giving it a good score? This study dives into that mystery, asking if a simple trick (a special prompt) actually makes medical AI safer, or if the safety score is just an illusion created by the judge's own biases.

The researchers in this paper set up a massive, high-stakes game of "spot the difference." They took four of the smartest medical AI models available (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, and Grok 4.3) and asked them 1,200 tricky medical questions. Half the time, they asked the models the questions normally. The other half, they gave the models a special "evidence-sufficiency wrapper"—think of it as a strict rulebook that forced the AI to list what evidence it had, what was missing, and to admit if it didn't have enough info to make a diagnosis.

The results were a mix of good news and a very important warning. When the primary AI judge (GPT-5.4-nano) graded the answers, the special rulebook worked wonders. It cut the number of "unsafe, overconfident" answers almost in half, dropping from 49.3% down to 24.7%. That's a huge drop! It meant the models were finally saying, "I don't know enough to be sure," instead of guessing.

However, the story gets twisty when you change the referee. The researchers brought in a different judge (Claude Sonnet 5) to grade the same answers. This new judge agreed that the models were safer, but it thought the improvement was only half as big. Instead of a 24.7 point drop, it only saw a 13.1 point drop. This proves that the "safety score" isn't a fixed, perfect number like a thermometer reading; it depends heavily on which AI is doing the judging. The primary judge was like a very sensitive smoke alarm: it almost never missed a real fire (it caught every unsafe answer), but it also went off for burnt toast (it flagged many safe answers as unsafe).

There was another catch: being safer came with a price tag. The "safety rulebook" made the models more cautious, but sometimes too cautious. When the models were asked questions they could actually answer, the rulebook made them hesitate and refuse to give a diagnosis much more often. For one model, Gemini 3.5 Flash, this was a disaster. Its ability to give the correct diagnosis crashed from 75% down to just 17%. It was so scared of being wrong that it stopped being helpful. Another model, GPT-5.5, handled the trade-off much better, losing almost no accuracy while gaining safety.

The researchers also made sure the models weren't just "faking it" by copying the special format the judge liked. They built a control group where the models used the same fancy headings but were told to give a definite answer anyway. The models still gave bad, unsafe answers in that group, proving that the safety gain came from the instruction to be careful, not just from using the right words.

In the end, this paper doesn't say we have a magic button that makes medical AI perfectly safe. Instead, it shows us that safety scores are relative. A "safe" AI for one judge might look risky to another. The best approach isn't to trust a single number, but to look at the direction of the change: the special prompt does make models act more carefully and ask for missing info, but it also makes them less helpful for some models. Before we let these robots into a real hospital, we need to know exactly which model we are using and which judge we are trusting, because the balance between being safe and being helpful is different for every single robot.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →