Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Policy Evaluation
This paper demonstrates that while chain-of-thought prompting enhances Large Language Models' performance on intuitive policy evaluations, their counterfactual reasoning remains significantly impaired on counter-intuitive cases due to an inability to fully override intuitive priors, revealing a critical dissociation between knowledge retrieval and logical reasoning.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: "Fast Thinking" vs. "Slow Thinking" in AI
Imagine you are trying to solve a puzzle. Sometimes the answer is obvious, like "If you drop a glass, it breaks." Other times, the answer is tricky and goes against your gut feeling, like "If you fine people for being late to pick up their kids, they actually become more late."
This paper tests whether Large Language Models (LLMs)—the smart AI chatbots we use today—are good at solving these puzzles, especially the tricky ones that go against common sense. The researchers used a special "test" based on real economic studies to see if the AI can think clearly when its "gut instinct" is wrong.
The Test: A Menu of 40 Real-World Scenarios
The researchers created a menu of 40 real-life policy stories drawn from actual economics and social science research. They sorted these stories into three categories based on how "obvious" the answer feels to a normal person:
- Obvious: The answer matches what you'd guess (e.g., "If you make healthcare cheaper, more people use it").
- Ambiguous: It's a toss-up; you could argue either way.
- Counter-Intuitive: The answer is the opposite of what you'd guess (e.g., "Fines for late daycare pickups make parents more late because they treat the fine as a price to pay for being late, rather than a moral obligation").
They asked four different top-tier AI models to solve these 40 stories using five different ways of asking (like just asking directly, pretending to be an expert, or asking the AI to "think step-by-step"). In total, they ran 8,000 experiments.
The Three Big Surprises
The results revealed three major findings that challenge how we think about AI reasoning:
1. The "Chain-of-Thought" Paradox
You might think that asking an AI to "think step-by-step" (a technique called Chain-of-Thought) would help it solve every hard problem.
- The Reality: It works great on Obvious problems. It's like giving a calculator to someone who already knows the answer; it just confirms it.
- The Catch: On Counter-Intuitive problems, the "step-by-step" help drops significantly.
- The Analogy: Imagine a person trying to walk through a maze. On a straight path (Obvious), telling them "look where you are going" helps them walk faster. But on a tricky path where the walls are moving (Counter-Intuitive), telling them to "look where you are going" doesn't help much because they keep walking in the wrong direction based on their first guess. The AI starts its "thinking" with the wrong idea, and even when it tries to think slowly, it can't fully shake off that first wrong guess.
2. The "Case" Matters More Than the "Brain"
You might think that a newer, bigger AI model would be much smarter than an older one.
- The Reality: The specific story being told mattered way more than which AI model was used. Some stories were just too hard for any of the models to solve correctly.
- The Analogy: Imagine a group of students taking a test. Whether the student is a genius or an average learner matters less than which question they are answering. Some questions are just so tricky that even the smartest student gets them wrong. In this study, the difficulty of the specific policy story explained about 67% of the mistakes, while the choice of AI model explained very little.
3. Knowing the Answer Doesn't Mean You Can Use It
The researchers checked if the AI got the right answer just because it had "read" about the topic before (like if a famous study was cited a lot).
- The Reality: There was no connection between how famous a study was and whether the AI got it right.
- The Analogy: Imagine a student who has memorized the entire textbook. If you ask them a simple question, they get it right. But if you ask a tricky question that requires them to apply the knowledge in a new way, they might fail, even if they've read the answer before. The AI "knows" the facts (it has the data), but when the facts contradict its gut feeling, it fails to use that knowledge correctly. It's like having a map but refusing to look at it because your feet want to go the other way.
The Conclusion: "Slow Talking" vs. "Slow Thinking"
The paper concludes that while these AI models are getting better at "thinking," they aren't quite there yet.
- System 1 (Fast Thinking): This is the AI's gut instinct. It jumps to the most common answer.
- System 2 (Slow Thinking): This is the AI's "step-by-step" reasoning.
The study shows that when the AI tries to use "Slow Thinking" on a tricky problem, it often just ends up "Slow Talking." It writes out a long, logical-sounding explanation, but the very first step of that explanation was based on a wrong gut instinct. Because it started with the wrong idea, the rest of the "slow thinking" just builds a fancy house on a shaky foundation.
The Takeaway: AI is great at confirming what we already expect, but it struggles to tell us when we are wrong, especially when the truth is surprising. If we use these tools for important decisions (like government policies), we need to be very careful, because the AI might confidently give us the wrong answer just because it feels "right" to its gut.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.