Absurd World: A Simple Yet Powerful Method to Absurdify the Real-world for Probing LLM Reasoning Capabilities
This paper introduces "Absurd World," a benchmarking framework that systematically alters real-world scenarios into logically coherent but absurd contexts to evaluate whether large language models possess robust logical reasoning capabilities independent of memorized real-world patterns.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart robot assistant that has read almost every book, article, and website on the internet. You ask it a simple question: "If I have 3 apples and eat 1, how many do I have left?" It answers "2" instantly. It seems perfect.
But what if you told the robot, "In this magical world, if you eat an apple, you actually gain an apple"? A human would instantly understand the new rule and say, "Okay, so I have 4 apples." But this paper suggests that many of our current AI assistants might get confused. They might ignore your new rule and just say "2" because that's what they've learned to expect in the "real world."
This paper introduces a new testing ground called Absurd World to see if AI can actually follow your rules, or if it's just blindly following its own memory of how the world usually works.
The Big Idea: The "Magic Soccer" Game
The researchers created a simple game based on a soccer penalty shootout. In a normal game, if you kick the ball into the net, you get a point. If you miss, you get nothing.
In Absurd World, they took this familiar game and twisted the rules just a little bit, creating "absurd" versions that are still easy for humans to understand but weird for AI:
- The "Miss" Rule: In this version, missing the net gives you a point, and hitting the net gives you zero.
- The "Least" Rule: The team with the lowest score wins.
- The "Ice Cream" Rule: Instead of points, teams earn ice cream cones.
- The "Switch" Rule: The teams shoot the net at the ball (swapping the roles of the objects).
The goal wasn't to make the math hard. The math was still just simple counting. The goal was to see if the AI could ignore its real-world training and strictly follow the weird, new instructions provided in the prompt.
How They Tested It
They acted like mad scientists, taking a standard soccer game and breaking it down into building blocks:
- Symbols: The players, the ball, the net.
- Actions: Hitting or missing.
- Rules: How points are scored and who wins.
They then swapped these blocks around to create hundreds of these "absurd" scenarios. They asked various AI models to read a short story about a match and declare the winner.
The Surprising Results
The paper found three main things that might make you rethink how "smart" these AIs really are:
1. The "Expensive" Models Were Actually Worse
You might think that the most expensive, powerful AI models would be the best at following new rules. Surprisingly, the "cheap" models often did better than the "expensive" non-reasoning models. The expensive models seemed to get stuck in their own heads, relying too heavily on what they thought should happen in a real soccer game, rather than what the prompt actually said.
2. The "Reasoning" Models Were the Champions
The models specifically designed to "think" (often called reasoning models) were the only ones that got a perfect score. They could look at the absurd rule ("Missing gives a point") and say, "Okay, I will ignore my real-world knowledge and follow this new rule." They didn't get confused by the magic.
3. Giving Examples Backfired
Usually, when you want to teach an AI something, you give it a few examples first (this is called "few-shot prompting"). You might say, "Here is an example of how to play, now you try."
- The Twist: In this study, giving the AI examples actually made it worse. It seemed that seeing examples of the "old" way of doing things confused the AI when it tried to switch to the "absurd" rules. It was like showing someone a picture of a cat before asking them to draw a dog; they kept drawing cat features.
The Takeaway
The paper argues that we need to stop just testing AI on hard, complex puzzles. Instead, we need to test them on simple tasks with weird rules.
If an AI can't follow a simple rule that says "Miss the goal to win," it might not be safe to trust it with more complex jobs where the rules are different from what it learned in its training data. Absurd World is a tool to check if an AI is truly listening to you, or if it's just reciting what it thinks you should hear.
In short: Just because an AI knows everything about the real world doesn't mean it can play by your rules in a magical one. This paper built a playground to find out which AIs can actually switch gears and which ones are stuck in neutral.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.