WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation
This paper introduces WorldSimProbe, a capability-based diagnostic framework that formalizes an "Observable Simulator Contract" to rigorously evaluate the fidelity of action-conditioned world models through controlled suites assessing motion induction, interaction grounding, and dynamics, revealing systematic failures in existing models that traditional visual or task-based metrics overlook.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to make a sandwich. You don't just want the robot to look like it's making a sandwich; you need it to actually pick up the bread, spread the peanut butter, and not accidentally throw the jar across the room. In the world of artificial intelligence, scientists are building "world models"—think of them as a robot's internal daydream machine. These models try to predict what will happen next if the robot moves its arm in a certain way. For a long time, researchers mostly checked if these daydreams looked pretty and realistic, like a high-quality movie. But there's a catch: a movie can look amazing while breaking the laws of physics. If a robot's brain thinks it can push a cup but the cup doesn't move, or if it thinks it grabbed a spoon when it actually just waved near it, the robot will fail in the real world. This paper tackles the tricky question of how to tell if a robot's "daydream" is actually a faithful simulation of reality, or just a convincing lie.
The authors of this paper, a team from various universities and research labs, realized that current tests for these AI models were too easy. They were mostly asking, "Did the robot finish the task?" or "Does the video look good?" But they wanted to ask harder questions: "If I wiggle the robot's arm just a tiny bit, does the simulation wiggle the same way?" or "If I tell the robot to tap a glass, does it actually make contact, or does it just pretend to?" To answer this, they created a new diagnostic tool called WorldSimProbe. Think of it as a rigorous "lie detector test" for robot simulators. Instead of just watching the final movie, they poke and prod the AI's predictions at every step of the way to see if the cause-and-effect chain holds up.
They put six different open-source AI models through their paces using over 18,000 test cases across three different robotic environments. The results were eye-opening. They found that while these AI models are great at generating smooth, plausible-looking videos, they often fail the "physics test." For instance, when the researchers gave the models slightly different instructions, many models didn't change their predictions enough, acting like they were stuck in a loop rather than reacting to the new input. Even worse, the models frequently hallucinated interactions. They would show a robot "grabbing" an object that was actually too far away, or making a heavy object float when it should have stayed put. The paper explicitly argues against the idea that a successful-looking video means the model is faithful; a model can get the right answer (the task is done) for the wrong reasons (the physics were fake).
The study didn't just find problems; it categorized them. They discovered that models struggle most with "interaction grounding"—knowing when an object is actually being touched versus just being near it—and "dynamics," which is how objects move and react after being hit or pushed. For example, the models were surprisingly bad at simulating a "shake" or a "tap," often getting the motion completely wrong. However, they did find that the models were better at preserving the style of movement if they were trained on human data, suggesting that copying human motion helps, but doesn't fix the underlying physics issues.
Ultimately, the paper suggests that we need to stop just grading these AI models on how pretty their videos are. The authors propose a new standard called the "Observable Simulator Contract," which demands that if you give a robot an action, the simulation must show the exact physical motion that action causes, and the environment must react only to that motion. Their tests show that current models are far from perfect at this, often failing in systematic ways that could lead to real-world robots crashing or dropping things. By using WorldSimProbe, researchers can now pinpoint exactly where a model is lying to us, whether it's failing to react to small changes, pretending to touch things it can't reach, or getting the physics of a drop all wrong. This isn't just about ranking which AI is "best"; it's about building a transparent way to fix the cracks in the foundation before we let these robots loose in our homes and factories.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.