Token Probability Beats Verbalized Condence for Error Detection in Small Open-Weight Language Models: A Medical and Psychiatric Benchmark Study
This pilot study demonstrates that for small open-weight language models (1.2B–9.2B parameters) applied to medical and psychiatric tasks, extracting internal token probabilities is a significantly more reliable method for error detection than relying on the models' verbally expressed confidence, particularly for the smallest models where verbalized confidence performs at chance levels.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of artificial intelligence, there is a growing desire to use powerful computer programs to help doctors and psychiatrists make difficult decisions. These programs, known as large language models, can read vast amounts of medical literature and answer complex questions about human health. However, a critical problem remains: how does a doctor know when to trust the computer? If the machine says a diagnosis is certain, but it is actually wrong, the consequences could be severe. For years, researchers have tried to solve this by asking the computer to simply tell them how sure it is. They ask the model to attach a percentage score to its answer, a verbal report of confidence. The hope was that if the computer said it was only fifty percent sure, a human could step in and double-check the work. But recent studies have shown that these self-reported scores are often unreliable, sometimes no better than a random guess.
This uncertainty is especially pressing for smaller, cheaper versions of these computer programs that can run on a single office computer rather than massive, expensive data centers. These smaller models are the ones most likely to be used in private clinics where patient data must stay secure and costs must be low. A new study by Kunal Dhanda at the Indian Institute of Technology Kharagpur investigates whether these smaller models can be trusted to spot their own mistakes. The research compares two different ways of measuring certainty. The first is the verbal confidence, where the model is asked to say how sure it is. The second is a hidden signal called token probability. To understand this, imagine the computer is writing an answer word by word. At every step, it calculates the mathematical likelihood of choosing the next specific letter or word. The token probability is simply the computer's internal calculation of how likely it was to pick the exact answer it eventually chose. This study asks a direct question: is this hidden mathematical calculation a better indicator of truth than the model's own spoken words?
The researchers tested four different small computer models, ranging in size from very compact to moderately powerful. They used these models to answer 1,600 medical and psychiatric questions drawn from standard exams. For every single question, the team recorded both the model's spoken confidence score and its internal token probability. They then checked the answers against the correct solutions to see which signal was better at predicting errors. The results were clear and consistent across all four models. In every case, the internal token probability was a far better predictor of whether an answer was right or wrong than the verbal confidence score. The difference was not small; for the smallest model, the internal signal was significantly more accurate, while the spoken confidence was so poor it was indistinguishable from a coin flip.
The study also looked at how these signals would work in a real-world safety system, where a doctor might decide to ignore the computer's answer if the confidence score was too low. When the researchers used the token probability to filter out the most uncertain answers, the accuracy of the remaining answers improved steadily for all four models. However, when they used the verbal confidence scores to filter answers, the results were flat or even harmful. For the smallest model, relying on the spoken confidence actually made the system less accurate, because the model was confidently wrong just as often as it was confidently right. This finding helps explain a strange behavior observed in earlier experiments where the smallest model performed worse when asked to admit when it did not know the answer. The researchers suggest that the model was not actually "knowing" it was unsure; it was simply reporting a confidence score that carried no real information, while the useful signal remained hidden inside its own calculations.
The implications of this work are practical and immediate for anyone deploying these tools in a medical setting. If a local computer system can reveal its internal probability scores, that data should be used to decide which answers need human review. Relying on the model to speak its confidence is not just less effective; for the smallest models, it creates a false sense of security. The study concludes that while verbal confidence might be passable for some larger, more advanced models, it is a dangerous tool for the smaller, privacy-focused systems that are becoming common in clinics. The most reliable way to detect errors in these systems is to look at the numbers the computer is already calculating but never saying out loud. This approach offers a genuine path to safer, more accurate medical AI without requiring expensive hardware or complex new software.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.