WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation
This paper introduces WBench, a comprehensive multi-turn benchmark designed to systematically evaluate interactive video world models across five key dimensions—video quality, setting adherence, interaction adherence, consistency, and physics compliance—using 289 test cases and validated automatic metrics to reveal that no current state-of-the-art model excels in all areas.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are building a video game engine, but instead of coding the rules, you are teaching an AI to "dream" a world into existence. You give it a starting picture and say, "Now, walk forward, jump over that rock, and suddenly a helicopter appears." The AI then generates the next few seconds of video based on your command.
The problem is: How do we know if the AI is actually good at this?
Currently, there are many different ways to test these AI "world models," but they are like trying to judge a car by only testing its brakes, or only its speed, or only its paint job. There is no single, unified test that checks everything at once.
Enter WBENCH. Think of WBENCH as the ultimate "Driver's License Exam" for these interactive video AIs.
The Big Idea: The "World Simulator" Test
The authors created a comprehensive test suite called WBENCH. Instead of just asking the AI to make a pretty video, they put it in a virtual driving school where it has to handle a multi-turn conversation with a user.
Here is how the test works, broken down into five simple categories:
- The Look (Video Quality): Does the video look smooth and pretty, or is it flickering and blurry? This is like checking if the car's paint is shiny and the tires are round.
- The Memory (Setting Adherence): If you told the AI, "It's a snowy mountain with a red robot," does it keep the snow and the red robot throughout the whole video? Or does the robot suddenly turn blue and the snow melt? This tests if the AI remembers the rules of the world it created.
- The Obedience (Interaction Adherence): If you say "Turn left," does the camera actually turn left? If you say "The robot jumps," does it jump? This is the "steering wheel" test. The paper found that some AIs are great at making pretty pictures but terrible at actually listening to your commands.
- The Stability (Consistency): If the camera spins around in a circle and comes back to the start, does the world look exactly the same as it did before? Or did the buildings shift, or did the robot disappear? This tests if the AI has a stable "memory" of the world's layout.
- The Physics (Physical Compliance): If a ball falls, does it fall down? If a car crashes, does it crumple? Or does the ball float up to the sky and the car pass through the wall like a ghost? This tests if the AI understands how the real world (or the fantasy world) works.
The Test Drive: What They Did
The researchers built a dataset of 289 different "scenarios" (like a snowy mountain, a busy city, or a fantasy castle). For each scenario, they created a sequence of 1,058 instructions (turns).
They tested 20 different AI models (including big names like Kling, Wan, Genie 3, and others). They asked these models to follow instructions like:
- "Walk forward."
- "Jump."
- "Switch the view from behind the character to the character's own eyes."
- "Make it rain."
The Results: No Perfect Driver
After running the tests, the paper found some surprising things:
- No "Super AI" exists yet: Just like no single car is the fastest, safest, and most fuel-efficient all at once, no single AI model passed every part of the test. Some were great at making pretty videos but terrible at following directions. Others were great at navigating but forgot what the world looked like after a few seconds.
- Navigation is tricky: Moving the camera (walking or turning) is actually a separate skill from making the video look good. An AI can make a beautiful movie but fail to move the camera when asked.
- The "Long Game" is hard: The longer the conversation goes on (the more turns you take), the more the AI starts to make mistakes. It's like a human trying to remember a complex story; after a while, they start mixing up the details.
- Open Source is catching up: Some of the models that are free for anyone to use performed just as well as, or even better than, the expensive, closed-source models on specific tasks.
The Bottom Line
WBENCH is a new, standardized ruler for measuring interactive video AIs. It proves that while these models are getting better at making movies, they still struggle to be reliable "world simulators" that can consistently follow instructions, remember the past, and obey the laws of physics over a long period of time.
The paper doesn't say these models are ready to replace human game designers or run self-driving cars yet. Instead, it says: "Here is exactly where they are failing, so developers know what to fix next." It's a diagnostic tool to help the next generation of AI become truly interactive.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.