WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models
This paper introduces WorldArena, a unified benchmark and the EWMScore metric designed to systematically evaluate embodied world models across both perceptual fidelity and functional utility in downstream decision-making tasks, revealing that high visual quality does not always guarantee strong task performance.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are training a professional chef. To see if they are ready for a real kitchen, you wouldn't just look at a high-definition photo of a delicious meal they made (that’s just "looking good"). You would actually watch them cook, see if they follow a recipe, check if they can handle a knife without cutting themselves, and see if they can actually produce a meal that tastes right.
Currently, scientists are trying to build "World Models"—AI "brains" for robots. These brains are supposed to "dream" or predict what will happen next when a robot moves (e.g., "If I move my arm left, the cup will tip over").
The problem? Most researchers are only checking if the AI's "dreams" look like pretty movies. They are checking the "photo quality," but they aren't checking if the "physics" actually make sense.
This paper introduces WorldArena, which is like a comprehensive "Final Exam" for these AI brains.
The Three Parts of the Exam
Instead of just looking at the pictures, WorldArena tests the AI in three specific ways:
1. The "Eye Test" (Perception Quality)
This is the most basic level. Does the video look sharp? Is the lighting realistic? Does the object stay the same shape, or does it morph into a blob halfway through? It’s like checking if a movie has high-definition graphics or if it looks like a blurry, glitchy mess.
2. The "Kitchen Test" (Functional Utility)
This is where the real challenge begins. The researchers test if the AI can actually do things. They use three metaphors:
- The Data Engine (The Practice Chef): Can the AI generate enough realistic "practice videos" to help a robot learn how to cook without needing a real kitchen?
- The Policy Evaluator (The Food Critic): If we show a robot a video of its own actions, can the AI accurately predict if the robot succeeded or failed? (Is the AI a reliable judge?)
- The Action Planner (The Head Chef): Can the AI actually plan a sequence of moves to complete a task, like "pick up the hammer and hit the block"?
3. The "Human Touch" (Human Evaluation)
Since math and code can sometimes miss the "vibe" of reality, they brought in real people to watch the videos and say, "Yes, that looks like a real robot," or "No, that looks like a hallucination."
The Big Discovery: The "Pretty vs. Practical" Gap
The most important finding in this paper is a bit of a wake-up call. The researchers found a massive gap between looking good and being useful.
Think of it like this: Imagine an AI that generates a stunning, 4K, cinematic video of a robot picking up a glass. It looks like a Hollywood movie! But, if you actually try to use that AI to guide a real robot, the robot fails. Why? Because in the "movie," the glass might pass through the robot's hand like a ghost, or the physics might be slightly "off."
The paper proves that just because an AI can make a beautiful video doesn't mean it understands how the physical world actually works.
Why does this matter?
By creating WorldArena, the researchers have built a standardized "arena" where all AI developers can compete. It moves the goalposts from "Make a pretty video" to "Make a brain that a robot can actually trust to navigate the real world."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.