What-If World: A Causal Benchmark for General World Models in Embodied Scenarios
This paper introduces "What-If World," a new benchmark comprising 319 prompt pairs that evaluate video generation models on their ability to produce physically consistent, causally divergent outcomes in response to single-variable changes, revealing that current state-of-the-art models largely fail to reliably simulate embodied scenarios for planning and control.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very talented artist who can paint incredibly realistic scenes. If you ask them to paint "a car braking gently," they do a great job. If you ask them to paint "a car braking hard," they also do a great job. Both paintings look perfect on their own.
But here is the catch: What-If World asks a different question. It asks: "If you change the instruction from 'gentle' to 'hard,' does the result actually change in the way physics says it should?"
In the real world, a hard brake stops a car much sooner than a gentle one. But in the world of current AI video generators, the artist might paint two almost identical stopping distances, even though the instructions were totally different. The AI is good at painting the look of a car, but it's bad at understanding the logic of how the car moves.
The Problem: The "Look-Alike" Trap
Current tests for AI video models are like grading a student's homework by looking at each page individually.
- Page 1: "Draw a car braking gently." (The drawing looks good. Pass.)
- Page 2: "Draw a car braking hard." (The drawing looks good. Pass.)
The teacher never compares the two pages side-by-side. So, the AI gets an A even if both drawings show the car stopping at the exact same spot, completely ignoring the difference in instructions. This is called the Contrastive Bottleneck. The AI can make things look real, but it can't make them act differently when you change the rules.
The Solution: What-If World
The researchers built a new test called What-If World. Instead of grading one video at a time, they force the AI to generate pairs of videos based on the same starting scene but with one tiny change in the instructions.
Think of it like a "Spot the Difference" game, but the AI has to create the differences itself based on physics.
The Setup:
- The Anchor: They take a real photo of a car or a robot arm right before it moves. This is the "frozen moment" where everything is the same for both videos.
- The Twist: They give the AI two prompts that are identical except for one physical detail.
- Prompt A: "The robot arm pushes the block gently."
- Prompt B: "The robot arm pushes the block hard."
- The Test: The AI must generate two videos. The researchers (using a super-smart AI judge) check:
- Did the robot actually push? (Adherence)
- Did the block move in a physically possible way? (Physics)
- Did the background stay the same? (Environment)
- Crucially: Did the "hard push" video show the block moving further or faster than the "gentle push" video? (Outcome)
The Six Rules of the Game
To make the test fair and scientific, the researchers organized the changes into six specific "rules" of physics, split into two categories:
1. The Environment (The Stage):
- Surface Friction: Is the road icy or dry? (Does the car slide?)
- Material/Medium: Is the object a sponge or a brick? (Does it squish or stay hard?)
- Obstacles: Is there a cone in the way? (Does the car hit it or go around?)
2. The Action (The Actor):
- Force/Degree: How hard is the push or brake?
- Spatial Alignment: Where exactly is the robot grabbing?
- Temporal Sequencing: Did the robot grab before or after it moved?
The Results: The AI is Still a Novice
The researchers tested 9 of the smartest video AI models in the world (both open-source and expensive closed-source ones). The results were sobering:
- The Ceiling: Even the best model only passed about 52% of the tests. That means it failed nearly half the time.
- The Gap: The open-source models were even worse, clustering around 28%.
- The Illusion: Most models scored high on "single video" tests (the videos looked great) but crashed on the "paired" tests (the videos didn't differ enough). They were faking the physics.
- Visual Bias: The AI did better when the change was obvious to the eye (like a car zooming fast vs. slow) but failed miserably when the change was subtle or invisible (like how slippery the road is).
The Big Takeaway
The paper concludes that while these AI models are amazing at rendering (making pretty pictures), they are not yet reliable simulators (understanding how the world works).
If you want to use an AI to plan a robot's moves or simulate a self-driving car, you need it to understand that "hard brake" means "shorter stop." Right now, the AI is just guessing what a hard brake looks like, not what it does. Until they can pass the What-If World test, we can't fully trust them to make decisions in the real world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.