WorldMark: A Unified Benchmark Suite for Interactive Video World Models
This paper introduces WorldMark, the first unified benchmark suite and online arena that enables fair, standardized comparison of interactive video world models by providing a common set of scenes, action sequences, and evaluation metrics across diverse systems.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a food critic trying to judge the best chefs in the world. But here's the catch: Chef A only lets you taste their food in their own private kitchen with their own ingredients. Chef B only lets you taste theirs in a different kitchen with different spices. Chef C refuses to let you taste anything unless you speak a language only they understand.
Because everyone is cooking in their own isolated world, you can never truly say, "Chef A is better than Chef B." You can only say, "Chef A's food is good in Chef A's kitchen."
This is exactly the problem with Interactive Video World Models (AI that creates video games or virtual worlds you can walk through). Right now, every AI model has its own rules, its own "kitchen," and its own way of taking orders. Comparing them fairly is impossible.
Enter WorldMark. Think of WorldMark as the first standardized "Taste-Test" arena for these AI world creators.
The Three Big Problems WorldMark Solves
1. The "Language Barrier" (Unified Action Mapping)
The Problem: One AI understands "WASD" (the keyboard keys for moving forward/back/left/right). Another only understands "Pose your hand like this." A third only understands "Say 'Go forward' in a sentence."
The WorldMark Solution: WorldMark acts like a universal translator.
- Imagine you have a single remote control with standard buttons: Move Forward, Move Back, Turn Left, Turn Right.
- WorldMark takes your simple press of "Forward" and instantly translates it into whatever specific language that specific AI speaks.
- Result: Now, every AI is trying to do the exact same task at the exact same time, using the exact same instructions. It's an apples-to-apples comparison.
2. The "Missing Playbook" (The Test Suite)
The Problem: Before, every AI was tested on its own favorite scenes (like a specific Minecraft castle or a specific city). If an AI was bad at forests but good at cities, it could just hide in the city tests.
The WorldMark Solution: WorldMark provides a standardized obstacle course with 500 different challenges.
- The Scenes: It includes 50 different starting images, ranging from realistic cities and nature spots to cartoonish or "stylized" art.
- The Views: You can test the AI from a "First-Person" view (like wearing a GoPro) or a "Third-Person" view (watching a character from behind, like in a video game).
- The Difficulty: The tests get harder.
- Easy: Walk forward for 20 seconds.
- Medium: Walk forward, then turn around.
- Hard: Walk in a complex patrol route for a full minute without the world falling apart.
3. The "Scorecard" (The Evaluation Toolkit)
The Problem: How do you grade the video? Is it pretty? Did it move when you told it to? Did the world make sense?
The WorldMark Solution: They created a three-part report card for every video generated:
- Visual Quality (The "Pretty" Score): Is the video blurry? Is the lighting weird? Does it look like a nice photo or a mess?
- Control Alignment (The "Obedience" Score): If you told the AI to turn right, did it actually turn right? Or did it just slide sideways?
- World Consistency (The "Logic" Score): This is the big one. If you walk around a corner and come back, is the chair still there? Did the character's face melt? Did the sky turn into water? This checks if the AI remembers the world it created.
What Did They Find? (The Plot Twist)
When they ran all the models through this standardized test, they discovered some surprising things that previous private tests missed:
- Pretty doesn't mean Smart: One model (YUME) made the most beautiful, artistic-looking frames. But if you walked around in its world, the logic fell apart. It was like a beautiful painting that changes every time you blink.
- Smart doesn't mean Pretty: Another model (Genie 3) kept the world perfectly consistent (the chair stayed the chair), but the images were a bit less "artistic" than the others.
- The "Third-Person" Nightmare: When the AI had to show a character from behind (like a video game), many models completely lost their minds. They couldn't keep the camera steady or the character in frame. It's like a camera operator who suddenly forgets how to hold a camera.
- Specialists aren't Generalists: An AI trained specifically on Minecraft couldn't handle a realistic city scene at all. It's like a chef who is amazing at baking cakes but burns toast.
Why This Matters
Before WorldMark, it was like watching a boxing match where the fighters wore blindfolds and fought in different rooms. We didn't know who was actually the strongest.
WorldMark removes the blindfolds, puts everyone in the same ring, and gives them the same opponent. Now, researchers can finally see which AI is truly building a stable, interactive world and which one is just hallucinating pretty pictures.
They even built a website (World Model Arena) where anyone can watch these AI models fight side-by-side in real-time, like a live leaderboard for the best virtual world creators.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.