Theory-Grounded Evaluation of Human-Like Fallacy Patterns in LLM Reasoning
This paper evaluates 38 language models using the Erotetic Theory of Reasoning to demonstrate that as model capability increases, their errors become more aligned with specific human-like fallacy patterns rather than random mistakes, and that premise order significantly influences fallacy production.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to be a detective. You give it a stack of clues (premises) and ask it to solve a mystery. If the robot gets the answer wrong, you want to know: Is it making a random mistake, or is it making the same kind of mistake a human would make?
This paper, presented at a major AI conference, asks exactly that question. The researchers wanted to see if Large Language Models (LLMs)—the brains behind chatbots like the one you're talking to—are just "stupid" in random ways, or if they have a specific, predictable "human-like" way of getting things wrong.
Here is the breakdown of their findings, using some everyday analogies.
1. The "Human Glitch" Theory
The researchers used a special theory called the Erotetic Theory of Reasoning (ETR). Think of this theory as a "blueprint for human brain fog."
- The Analogy: Imagine your brain is a filter. When you hear a complex story, your brain tries to keep all the possibilities open. But sometimes, to save energy, your brain accidentally drops a possibility too early. This leads to a specific type of logical error called a "fallacy."
- The Tool: The team built a software tool called PyETR (a calculator for this theory). It generates thousands of logic puzzles designed specifically to trigger these "brain fog" moments in humans.
2. The Experiment: 38 Detectives, 383 Puzzles
They took 38 different AI models (ranging from small, cheap ones to massive, super-smart ones) and gave them 383 of these logic puzzles.
They looked for two things:
- Did the AI get the answer right?
- If it got it wrong, did it make the specific mistake that the "Human Brain Fog Blueprint" predicted?
3. The Big Surprise: Smarter AI = More Human-Like Mistakes
This is the most counter-intuitive finding. Usually, we think: "The bigger and smarter the AI, the fewer mistakes it makes."
The Reality:
- Overall Accuracy: As the AI got "smarter" (measured by how well it performs in general chatbot tournaments), its ability to get the right answer on these logic puzzles did not improve. It stayed roughly the same.
- The "Human" Error Rate: However, when the smarter AIs did get it wrong, they were much more likely to make the specific, human-like mistakes predicted by the theory.
The Metaphor:
Imagine a group of students taking a tricky math test.
- The "Dumb" Student: Gets the answer wrong because they can't do the math at all. They guess randomly.
- The "Smart" Student: Knows the math, but when they slip up, they slip up in a very specific way because they are overthinking it or relying on a bad habit.
- The Finding: As the AI models got "smarter" (more like the Smart Student), their errors stopped being random noise and started looking exactly like the specific, predictable errors humans make. They weren't just failing; they were failing like us.
4. The "Order of Operations" Trick
The researchers also played a trick on the AIs. In logic, the order of clues shouldn't matter. (If A implies B, and B implies C, it doesn't matter if you say A then B, or B then A).
But humans are weird. If you tell a human a story in a different order, they might get confused and make a mistake.
- The Test: The researchers gave the AIs the same logic puzzles but swapped the order of the clues.
- The Result: For many models, simply reversing the order of the sentences stopped them from making the human-like mistake. It was like giving the AI a "reset button" that broke the bad habit. This proves the AI isn't just calculating; it's reacting to the flow of the story, just like a human does.
5. Why Does This Matter?
This is a double-edged sword.
- The Good News: It means AI is becoming sophisticated enough to understand the nuances of human language and reasoning patterns. It's not just a calculator; it's starting to think like a person.
- The Bad News: It means that as AI gets more powerful, it might inherit our worst cognitive biases. If we use AI for medical diagnoses or legal advice, we can't just assume it will be "more logical" than us. It might be more logical in some ways, but it will still fall into the same traps our brains do.
The Bottom Line
The paper concludes that scaling up AI (making it bigger and smarter) doesn't automatically fix its reasoning. Instead, it seems to tune the AI to make the same systematic errors humans make.
The Takeaway: If you want to build a truly reliable AI, you can't just make it bigger. You have to specifically train it to avoid the "human glitches" that even our smartest models are now picking up. We need to teach the robot not just how to think, but how to stop thinking like a human when logic is required.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.