"very likely" Means "uncertain"? How LLMs Diverge from Humans in Linguistic Uncertainty Quantification
This paper investigates the divergence between humans and large language models in verbal uncertainty quantification by introducing an optimization-based algorithm, METHODNAME, which learns optimal probability mappings for verbal markers directly from model outputs to reveal systematic confidence disparities without relying on repeated sampling.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
When we speak, we often hedge our bets. We say something is "likely," "possible," or "almost certain," using these words to signal how much we truly know. This ability to express doubt is a form of metacognition, a mental check where we acknowledge the boundaries of our own knowledge. It is a crucial part of human communication, allowing us to navigate uncertainty without pretending to have answers we do not possess. In the world of artificial intelligence, specifically with large language models, this same ability is becoming a critical safety feature. These models can sometimes generate confident-sounding but factually incorrect information, a phenomenon known as hallucination. To trust these systems, especially in sensitive fields like medicine or law, we need to know when they are unsure. The challenge lies in figuring out if the words an AI uses to express doubt mean the same thing to a human listener as they do to the machine itself.
A team of researchers set out to investigate whether the verbal uncertainty expressed by large language models aligns with human understanding. They began by gathering a massive collection of how people interpret phrases like "very likely" or "uncertain." Drawing from decades of psychological and decision-science research, they compiled a reference guide that maps these common phrases to specific numerical probabilities. For instance, in the human mind, the phrase "very likely" corresponds to a high degree of confidence, while "possible" sits somewhere in the middle. The researchers then tested several advanced language models, asking them to answer questions and explain their confidence using these same natural language phrases. They compared the models' answers against the human reference guide to see if the machines were speaking the same language as their users.
The results revealed a significant disconnect. While the models could use the words, they assigned different levels of certainty to them than humans do. For example, when a human hears "very likely," they interpret it as a strong probability, but the models in the study often used this phrase to describe situations where they were actually quite unsure. Conversely, the models might express a high degree of confidence for phrases that humans consider to be mere guesses. This mismatch means that simply reading a model's output and assuming its verbal cues reflect its internal state can be misleading. The study showed that relying on a human-made dictionary of uncertainty words was not enough to accurately gauge a machine's confidence; the models had their own, distinct internal logic for what those words meant.
To bridge this gap, the researchers developed a new method called VOCAL. Instead of forcing the models to conform to human definitions, this approach learns the specific meaning of uncertainty words directly from the model's own behavior. The system analyzes thousands of the model's responses, noting which phrases it uses when it is right and which it uses when it is wrong. It then builds a custom map that translates the model's specific verbal habits into accurate probability scores. This process involves a mathematical optimization that smooths out the data, ensuring that similar-sounding phrases receive similar scores, even if the model uses them rarely. By tailoring the interpretation to the specific machine, the method creates a much more reliable way to detect when the model is unsure.
The experiments demonstrated that this customized approach works significantly better than trying to force a human standard onto the machine. In tests across various datasets, including complex math problems and medical questions, the new method was far more accurate at predicting whether a model's answer was correct or incorrect compared to using the human reference guide alone. In some cases, the improvement was substantial, turning a method that barely performed better than random guessing into one that could reliably spot errors. The researchers found that this technique could achieve performance levels comparable to much more expensive methods that require the model to generate multiple answers and compare them, but without the extra time or computing power.
The study concludes that while large language models can mimic human speech patterns, their internal understanding of uncertainty is fundamentally different from our own. A phrase like "very likely" does not carry the same weight for the machine as it does for a person. To make these systems truly trustworthy, we cannot simply ask them to speak like humans and expect them to mean the same thing. Instead, we must learn to read their specific dialect of uncertainty. By creating custom maps that translate a model's unique verbal signals into clear signals of confidence, we can better distinguish between a correct answer and a confident hallucination, making the technology safer and more reliable for real-world use.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.