The Riddle Riddle: Testing Flexible Reasoning in Large Language Models and Humans
This paper introduces the "riddle riddle" paradigm to demonstrate that while humans flexibly adapt their reasoning strategies to problem content, large language models often rely on pattern matching and surface features, leading them to inappropriately apply inventive reasoning to literal problems and suggesting their strong performance on genuine riddles may stem from memory retrieval rather than flexible reasoning.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are playing a game where you have to solve puzzles. Some puzzles are tricky riddles that require you to think outside the box, while others look exactly like riddles but are actually just simple math problems in disguise.
This paper introduces a new way to test how smart AI (Large Language Models) really are compared to humans. They call this the "Riddle Riddle" paradigm.
Here is the simple breakdown of what they did and what they found:
The Setup: The "Trick" vs. The "Literal"
The researchers created two types of questions that look almost identical on the surface:
- The Real Riddle (The Trick):
- Example: "A cowboy rides into town on Friday, stays for three days, and rides out on Friday. How is this possible?"
- The Answer: You have to realize "Friday" is the name of his horse, not the day of the week. This requires inventive thinking.
- The "Riddle Riddle" (The Literal):
- Example: "A cowboy rides into town on Friday, stays for three days, and rides out on Monday. How is this possible?"
- The Answer: This is just simple math. Friday + 3 days = Monday. No tricks needed. This requires literal thinking.
The key is that both questions look the same. They have the same sentence structure and rhythm. The only difference is whether a "trick" is actually needed to solve them.
The Experiment
The researchers asked two groups to solve these:
- Group 1: Nine of the smartest AI models available today (like GPT-5, Claude, Gemini, etc.).
- Group 2: 100 human adults.
The Results: A Perfect Reversal
The results were like looking in a mirror, but the reflection was flipped upside down.
1. How the AI Acted: The "Look-First" Solver
The AI models were super good at the Real Riddles (85% correct) but struggled with the Riddle Riddles (only 50% correct).
- What happened: When the AI saw a question that looked like a riddle, it immediately assumed, "Ah, this must be a trick! I need to use my creative, outside-the-box thinking."
- The Mistake: Even when the question was a simple math problem (the Riddle Riddle), the AI still tried to find a hidden trick. It was like a detective who sees a locked door and immediately assumes there's a secret tunnel, even when the door is just unlocked and open. The AI was so used to riddles having tricks that it couldn't stop looking for them.
2. How Humans Acted: The "Literal-First" Solver
Humans showed the exact opposite pattern. They were great at the Riddle Riddles (80% correct) but struggled with the Real Riddles (50% correct).
- What happened: Humans naturally default to taking things literally. When they saw the "Friday + 3 days = Monday" question, they solved it instantly.
- The Mistake: When faced with the Real Riddle (the horse named Friday), humans often got stuck thinking, "Wait, Friday is a day of the week, so this is impossible!" They had to work harder to override their literal thinking to find the trick.
The Big Takeaway
The paper argues that AI isn't actually "reasoning" flexibly yet.
- Humans are flexible but lazy. They prefer simple, literal answers and only switch to "creative mode" if they realize they are stuck.
- AI is the opposite. It has memorized that "Riddle-like structure = Creative Answer." It doesn't actually check if a creative answer is needed; it just sees the shape of the question and automatically pulls out the "creative toolbox."
The authors call this the "Illusion of Reasoning." Just like a person might see a visual illusion in a picture even when the lines are actually different lengths, the AI "sees" a riddle trick even when the problem is just a simple math question.
The Conclusion
The paper concludes that while AI can get the right answer on famous riddles, it might just be reciting from memory or following a rigid rule ("If it looks like a riddle, use a trick"). It hasn't yet learned the human skill of pausing to ask, "Does this problem actually need a trick, or can I just solve it normally?"
In short: AI thinks outside the box whenever the box has a label that says "Riddle," even if the box is empty. Humans usually stay inside the box, only looking outside when they have to.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.