Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?
This paper introduces Target-Bench, a novel benchmark comprising 450 robot-collected scenarios designed to evaluate video world models' semantic reasoning and mapless path planning capabilities, revealing a significant gap between visual realism and planning performance while demonstrating that fine-tuning on real-world data can substantially improve task-level results.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very talented artist who can paint incredibly realistic pictures of the future. If you show them a photo of a living room and say, "Draw what happens next," they might paint a cat jumping onto a sofa, a ball rolling across the floor, and a person walking through the door. The picture looks perfect, the lighting is right, and the physics seem real.
But here's the catch: If you asked that artist to actually walk through the room without bumping into the furniture, could they do it?
This is the exact problem researchers are tackling with a new project called Target-Bench.
The Big Idea: From "Watching" to "Walking"
For a long time, AI video models (like the ones that make deepfakes or generate movie scenes) have gotten amazing at looking real. They can predict how a scene evolves visually. But the robotics world wants to use these models to help robots act.
The researchers asked: If an AI can predict a video of a robot walking toward a specific object (like a red chair), can we actually use that video to tell a real robot how to move?
To find out, they built Target-Bench, which is like a driving test for AI video models.
The Test Drive: Target-Bench
Think of Target-Bench as a giant obstacle course with 450 different scenarios.
- The Setup: They used a real robot dog (a quadruped) to walk through real houses and parks.
- The Goal: The robot was told to go to specific things, like "the blue car," "the open door," or even tricky things like "the place where you'd sit to rest" (an implicit goal).
- The Challenge: The researchers took these real videos and fed them into various AI video models. The models tried to guess what the next few seconds of the video would look like if the robot kept moving toward that goal.
The "Decoder" Translator
Here is the clever part. The AI models just output a video file. A robot can't drive on a video file; it needs coordinates (turn left, go forward 2 meters).
The researchers built a special tool called a World-Decoder. Imagine this as a translator that watches the AI's generated video and says, "Okay, in this frame, the camera moved 1 meter forward. In the next, it turned slightly right." It turns the visual prediction back into a physical path.
The Scorecard: It's Not About Perfect Lines
In the past, people judged these models by asking, "Does the generated video look exactly like the real video?"
Target-Bench changes the rules. It asks: "Did the robot head in the right direction?"
They use a scoring system that cares more about tendency than perfection.
- The Analogy: Imagine you are trying to walk to a coffee shop.
- Old Way: Did you walk on the exact same sidewalk tiles as the person before you? (Too strict!)
- Target-Bench Way: Did you generally walk toward the coffee shop without walking into a wall or turning around? (Practical!)
They measure things like:
- Did you get close? (Endpoint Score)
- Did you stay on the right path? (Approach Consistency)
- Did you wander off too far? (Miss Rate)
The Results: The "Uncanny Valley" of Planning
The results were a bit of a reality check.
- The Visuals: The best AI models (like Sora 2 or Wan2.2) generated beautiful, realistic videos.
- The Planning: When they tried to turn those videos into a path for a robot, the scores were surprisingly low. The best model only got a 34% score.
What does this mean? It means the AI is great at imagining the future, but it's still bad at planning for it. It's like a movie director who can visualize a perfect chase scene but doesn't know how to drive a car.
The Good News: A Little Training Goes a Long Way
The researchers then tried something cool. They took one of the open-source AI models and gave it a "crash course" using just a small amount of real robot data (325 videos).
The Result: After this tiny bit of training, the model's planning score jumped from 8% to nearly 40%.
The Metaphor: It's like taking a person who has only watched thousands of driving movies and giving them 10 hours of actual driving lessons. Suddenly, they aren't just a movie critic; they can actually drive the car.
Why This Matters
This paper is a wake-up call and a roadmap.
- Wake-up Call: Just because an AI can make a pretty video doesn't mean it understands how to navigate the real world.
- Roadmap: We don't need massive supercomputers to fix this. A little bit of real-world data can teach these "movie makers" how to become "drivers."
In the future, this could mean robots that can look at a messy room, understand where the "clean spot" is, and figure out how to walk there without a pre-made map, just by "imagining" the path first.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.