Humans and LLMs Diverge on Probabilistic Inferences
This paper introduces ProbCOPA, a dataset of probabilistic inferences annotated with human judgments, to demonstrate that state-of-the-art reasoning LLMs consistently fail to replicate the graded and varied nature of human probabilistic reasoning, highlighting a critical gap in evaluating models beyond deterministic tasks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: The "Maybe" Machine vs. The "Maybe" Human
Imagine you are standing on a highway watching a car accident.
- Human thought: "Oh no, an accident! Traffic is probably going to be a nightmare." (You feel pretty sure, but you know it's possible the police cleared it instantly, or everyone took a different route. So, you think it's likely, but not 100% guaranteed.)
- LLM thought: The paper argues that current AI models struggle with this "maybe." They tend to act like a robot that only sees black and white. They either think "Traffic will definitely be bad" or "Traffic will definitely be fine," often skipping the messy, uncertain middle ground where real life happens.
The researchers built a new test called PROBCOPA to see how well AI handles these "gray area" guesses compared to real humans.
1. The Test: A New Kind of Riddle
The researchers took a classic puzzle dataset called COPA (which asks things like, "It rained. Did the grass get wet?") and turned it into a probability test.
Instead of asking "True or False?", they asked: "On a scale of 0 to 100, how likely is this outcome?"
- 0 = Impossible.
- 100 = Guaranteed.
- 50 = A coin flip.
They asked 30 different humans to rate 210 different scenarios. Then, they asked 8 of the smartest AI models (like GPT-5, Gemini, Claude) to do the exact same thing, 30 times each.
2. What Humans Did: The "Goldilocks" Zone
When humans looked at these scenarios, their answers were graded and varied.
- The Analogy: Imagine humans are a choir singing a chord. Some sing a little higher, some a little lower, but they all cluster around a specific note.
- The Result: Humans agreed on the extremes (0 or 100). But for the middle stuff, they gave a wide range of answers. Some said "70% likely," others said "40% likely." This spread of answers is actually a good thing—it shows humans understand that the world is uncertain.
3. What the AI Did: The "All-or-Nothing" Robot
The AI models behaved very differently.
- The Analogy: Imagine the AI is a light switch. It's either ON (100% sure) or OFF (0% sure). It rarely leaves the light dim.
- The Result: The AI almost never gave answers in the middle (like 40% or 60%). It jumped straight to "Highly Likely" or "Highly Unlikely."
- The Problem: Even when humans were confused and gave different answers, the AI gave the same answer every single time. It lacked the "human variation" that comes from genuine uncertainty.
4. The "Thinking" Process: Why does the AI do this?
The researchers peeked inside the AI's "brain" (its reasoning chain) to see how it thought.
- The Analogy: When a human is unsure, they might say, "Well, it could be this, or maybe that... I'm not sure."
- The AI's Trick: The AI does mention alternatives. It says, "Maybe the traffic cleared, maybe it didn't." But then, instead of saying, "I'm torn," it immediately picks a side and acts 100% confident in that choice.
- The Finding: The AI spends more time "thinking" (generating more text) when humans are confused, but it still doesn't produce a range of answers. It just thinks harder about the wrong thing: how to be decisive, rather than how to be uncertain.
5. The "Temperature" Experiment: Can we fix it?
The researchers tried to force the AI to be more "human-like" by changing its settings (like turning up the "temperature" to make it more random, or giving it a fake persona like "You are a skeptical detective").
- The Result: It didn't work.
- Turning up the "randomness" just made the AI start hallucinating or repeating nonsense.
- Giving it a "persona" didn't make it vary its answers enough.
- Making it "think longer" didn't help either.
The Takeaway: We Need "Gray" Thinking
The paper concludes that while AI is amazing at math and logic (where the answer is 100% right or 100% wrong), it is terrible at probabilistic reasoning (where the answer is "probably, but maybe not").
- Human Reasoning: "I think there's a 60% chance of rain, but I'll bring an umbrella just in case."
- AI Reasoning: "It is raining." (Even if the data says it's only 60% likely).
Why does this matter?
As we start using AI for things like medical diagnoses, legal advice, or self-driving cars, we need machines that understand uncertainty. If an AI says a surgery has a "99% success rate" when it's actually a coin toss, that's dangerous. This paper shows that right now, our best AI models are still too confident and too rigid to truly understand the messy, uncertain nature of human life.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.