Fixing Abstention with Token Probability: A Budget-Matched Causal Test Across Four Small Open-Weight Models
This paper demonstrates that replacing verbalized confidence with a budget-matched token-probability abstention policy causally eliminates a specific accuracy-degrading effect observed in the llama3.2:1b model, but fails to generalize this benefit to other small open-weight models, indicating that such fixes are model-specific and that fixed-budget abstention can remain net-harmful even with improved targeting.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the rapidly evolving world of artificial intelligence, researchers are constantly teaching computer programs to answer questions. But a critical challenge remains: knowing when a program does not know the answer. When a human is unsure, they can say, "I don't know." This act of stepping back, called abstention, is often seen as a safety feature, preventing the machine from guessing wildly and spreading misinformation. However, recent investigations have revealed a strange glitch in some of these systems. When asked to admit uncertainty, the smallest and most lightweight models in a specific family of open-source programs sometimes perform worse than if they had simply been forced to guess every time. It appears that asking these models to speak about their own confidence triggers a breakdown, causing them to lose accuracy on the very questions they were trying to handle carefully.
This new study takes a closer look at that breakdown to see if it can be fixed. The researchers focused on four small, open-weight models—versions of artificial intelligence that are compact enough to run on standard computers but still capable of complex reasoning. They had previously observed that when the smallest model was prompted to say "I don't know" if it felt unsure, its accuracy dropped significantly across medical and psychiatric questions. The team suspected the problem was not the act of abstaining itself, but rather the method used to decide when to abstain. The model was being asked to verbalize its confidence, a signal that turned out to be statistically useless, essentially no better than a coin flip. Meanwhile, the internal math the model uses to generate each word, known as token probability, was a much stronger indicator of whether an answer was correct. The question was whether switching the decision-making process from the model's spoken confidence to its internal math would stop the accuracy drop.
To find the answer, the researchers did not run any new experiments or ask the models to generate new text. Instead, they performed a careful re-analysis of data already collected in three previous studies. They took the existing records of how the models answered questions and compared two different ways of deciding when to skip an answer. The first method was the original one: the model would skip an answer only if it verbally stated it was unsure. The second method was a new policy: the model would skip an answer if its internal math showed a low probability of being correct. Crucially, the researchers adjusted the second method so that it skipped the exact same number of answers as the first method. This ensured that any difference in results was due to which questions were skipped, not how many.
The results were striking for the smallest model. When the researchers replaced the verbal confidence check with the internal probability check, the severe drop in accuracy vanished. Under the old method, the model's accuracy had fallen by seven percentage points, a significant loss. Under the new method, the drop shrank to nearly zero, a change that was statistically indistinguishable from no effect at all. This fix worked consistently across both general medical questions and psychiatric scenarios. It confirmed that the original failure was indeed caused by the model's inability to accurately report its own confidence, and that using its internal math as a guide successfully corrected the error.
However, the study also revealed that this solution is not a universal cure. When the researchers applied the same fix to a different model in the group, one that had also suffered from the verbal confidence problem, the results were different. This second model had the best internal math signal of all the models tested, yet the accuracy still dropped significantly when the new policy was applied. The researchers found that this model had a much higher rate of saying "I don't know" in the first place. Because it skipped so many questions, it was discarding a large number of correct answers along with the wrong ones. Even though the internal math was good at sorting the answers, the sheer volume of skipped items was too high to be beneficial. This suggests that for models that are already very accurate, simply changing the signal used to decide what to skip is not enough; the number of items skipped may need to be reduced instead.
The study concludes that while routing the decision to skip through internal math rather than spoken confidence is a powerful fix for specific failures, it is not a one-size-fits-all solution. For the smallest model, the change was a complete rescue, turning a harmful instruction into a neutral one. For the larger, more accurate model, the problem was not the signal but the budget of skipped items. The findings emphasize that in artificial intelligence, there is no single safety patch that works for every system. Each model requires its own specific validation to understand how it handles uncertainty, and what works as a correction for one may not work for another. The researchers caution that these results are based on a specific set of data and should be viewed as hypotheses for further study rather than final instructions for deployment, but they offer a clear, causal proof that fixing the channel of communication can sometimes fix the outcome.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.