FALSIFYBENCH: Evaluating Inductive Reasoning in LLMs with Rule Discovery Games
This paper introduces FALSIFYBENCH, a benchmark inspired by the Wason 2-4-6 task to evaluate LLMs' inductive reasoning for scientific discovery, revealing that reasoning models outperform instruction-tuned ones primarily due to their capacity for negative testing and that failure stems from identifiable patterns in hypothesis navigation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are playing a guessing game with a friend who knows a secret rule, but you don't. Your friend gives you three examples that fit the rule, like "2, 4, 6." Your job is to figure out the secret rule by proposing your own sets of numbers and asking, "Does this fit?"
This is the classic "Wason 2-4-6 task," a famous puzzle used to study how humans think. The paper you're asking about, FALSIFYBENCH, takes this game and gives it to Artificial Intelligence (AI) models, but instead of numbers, they use words and categories (like "animals" or "tools").
Here is the breakdown of what the researchers did and found, using simple analogies.
The Game: "Find the Hidden Category"
Think of the AI as a detective trying to solve a mystery.
- The Setup: The AI is shown three items, say, a "bat," a "skimmer," and a "tarsier."
- The Goal: The AI must guess the hidden category these belong to (e.g., "Animals").
- The Catch: The AI doesn't know the answer. It has to guess a rule, test it, and get feedback.
- If the AI guesses "Flying things" and tests a "bat," the system says, "Yes, that fits."
- If the AI tests a "fish" and the system says, "No, that doesn't fit," the AI learns its rule is too broad.
- If the AI tests a "whale" and the system says, "Yes, that fits," but the AI thought "Flying things," the AI learns its rule is too narrow.
The Big Problem: The "Yes-Man" Trap
The researchers discovered that most AI models (and humans) suffer from a specific thinking error called Confirmation Bias.
Imagine you are a detective who is convinced the culprit is "The Butler."
- The "Yes-Man" Strategy (Confirmation): You only ask questions that prove the Butler did it. "Did the Butler have a motive?" "Yes." "Was the Butler in the room?" "Yes." You feel confident, but you never actually check if the Butler is innocent. You are just looking for "Yes" answers.
- The "Devil's Advocate" Strategy (Falsification): You actively try to prove the Butler is not the culprit. You ask, "Was the Butler at the beach at the time of the crime?" If the answer is "Yes," your theory is destroyed, and you have to find a new suspect.
The Paper's Main Finding:
The AI models that acted like "Devil's Advocates" (trying to break their own theories) were much better at solving the puzzle. The models that acted like "Yes-Men" (only looking for proof that they were right) got stuck. They kept guessing narrow rules (like "flying mammals") and never realized the real rule was broader ("all animals").
The Results: Who Played Best?
The team tested 12 different AI models. Here is what they found:
- The "Reasoning" Models Won: The newer AI models designed to "think" before they speak (often called "reasoning models") were better detectives than the standard models. They were more likely to try to break their own theories.
- No One is Perfect: Even the best AI models didn't solve the game perfectly. They still made mistakes, showing that AI is not yet a master of scientific discovery.
- It's Not About Being Smart, It's About Strategy: The researchers checked if the AI was failing because it didn't know the definitions of words (like not knowing what a "bat" is). They found that wasn't the problem. The AI knew the words; it just had a bad strategy. It was too afraid to test ideas that might be wrong.
The "Turn-by-Turn" Analysis
The researchers didn't just look at who won; they watched how the AI played every single move. They found two main ways the AI failed:
- Getting Lost in the Woods: Instead of slowly moving from a specific guess to a general one (like going from "bats" to "mammals" to "animals"), the AI would sometimes jump to completely unrelated guesses (like "things that start with the letter B").
- Focusing on the Wrong Details: Sometimes, instead of guessing the meaning of the words, the AI would guess based on how the words looked. For example, it might guess, "The rule is that all the words have three letters," or "The rule is that they all start with a vowel." This is like a detective ignoring the crime scene and focusing on the color of the suspect's shoes.
The Bottom Line
This paper introduces a new test (FALSIFYBENCH) to see if AI can do real scientific thinking. The conclusion is that while AI is getting better at "thinking," it still struggles with the most important part of science: the courage to be wrong.
To be a good scientist (or a good detective), you have to actively try to prove your own ideas false. The AI models that learned to do this were the winners, but even they have a long way to go before they can replace human scientists.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.