Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language
This study evaluates nine large language models as post-hoc explainers for probabilistic risk communication and finds that while they are generally consistent, they remain significantly miscalibrated—particularly regarding uncertainty—indicating they are not yet reliable zero-shot tools for translating numerical predictions into natural language.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a crystal ball that can predict the future, but instead of giving you a clear "Yes" or "No," it whispers numbers like "73% chance" or "very shaky prediction." Now, imagine you ask a super-smart robot (a Large Language Model, or LLM) to translate those numbers into plain English for a regular person. You want the robot to say, "It's very likely to rain," when the number is high, and "It's very uncertain," when the number is shaky.
The big question this paper asks is: Can these robots actually do that job without messing up?
The researchers set up a massive game of "translation" to find out. They didn't just ask the robots to guess; they fed them simulated predictions from a "crystal ball" (mathematically modeled using something called a Beta distribution) that had specific, known levels of likelihood and uncertainty. They then asked nine different AI models to pick the perfect phrase from a strict menu of options (like "very likely" or "highly uncertain") to describe the situation. They did this over and over again, changing the "temperature" (how creative or random the robot is allowed to be) and testing them in six different real-world scenarios, from predicting floods to guessing if a gambler will win.
Here is the verdict, straight from the lab: The robots are consistent, but they are terrible at calibration.
Think of it like a weather forecaster who is incredibly reliable at saying the same thing every time you ask, but is completely wrong about what that thing means.
- Consistency (The "Reliable Robot"): If you ask the same robot the same question ten times, it will almost always give you the exact same answer. It's like a parrot that never forgets its script. Most of the models scored very high here, with one top-tier model (GPT-5.4) getting a perfect score of 1.0. It never wavered.
- Calibration (The "Truth-Teller"): This is where the robots stumble. Calibration means matching the words to the actual math. If the math says "95% chance," the robot should say "virtually certain." If it says "10% chance," it should say "very unlikely." The study found that most robots are miscalibrated. They often pick the same middle-of-the-road phrases (like "somewhat certain") regardless of whether the math is screaming "almost impossible" or "guaranteed." They are like a person who always says "maybe" no matter what the dice roll is.
What the paper explicitly rules out:
You might think, "Maybe the robots just can't do the math! Maybe if we give them the final numbers directly instead of raw data, they'll get it right." The researchers tested this exact idea. They gave the robots the pre-calculated "mode" (the most likely number) and the "sample size" (how sure the math is) right in the prompt. The result? It didn't help. The robots still failed to map the numbers to the right words. This suggests the problem isn't that they can't do the math; the problem is in the translation step itself. They just don't have a solid internal dictionary that links numbers to words.
The "Temperature" Trade-off:
The researchers also played with the "temperature" setting, which controls how random the robot's choices are.
- Low Temperature: The robot is a boring robot. It picks the same safe word every time. High consistency, but it gets stuck in a rut and doesn't match the numbers well.
- High Temperature: The robot gets a bit wild. It starts picking different words, which actually helps it match the numbers better (improving calibration), but it becomes a flaky robot that gives you different answers for the same question (destroying consistency).
The One Star Performer:
There was one exception: GPT-5.4. This model was a beast at describing likelihood. It got a near-perfect calibration score of 0.96 and a perfect consistency score of 1.0. It seemed to have learned a perfect, rigid rulebook: "If the number is X, I say Y." However, even this super-robot failed when it came to uncertainty. It scored only 0.37 on uncertainty tasks, defaulting to vague, middle-ground phrases just like the others. It turns out, even the smartest robot doesn't know how to say "I'm really not sure" in a way that matches the math.
The Bottom Line:
The paper concludes that right now, these AI models are not ready to be standalone tools for explaining risk to the public. They are like a very consistent translator who speaks a language where the words don't quite match the meaning. If you ask them to explain a medical risk or a flood warning, they might give you a confident-sounding answer that is actually misleading.
The authors suggest that while these models are great at making fluent, easy-to-read sentences, we can't just plug them into a risk calculator and expect them to be accurate. The "bottleneck" isn't the math; it's the robot's inability to truly understand how to turn a number into a trustworthy word. Until we fix that, we need to be very careful about letting AI explain the odds of life-or-death situations.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.