Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models
This paper presents a systematic comparison of Vision-Language Models (VLMs) and Video Generation Models (VGMs) for spatial intelligence, revealing their complementary strengths in semantic understanding versus geometric reasoning and demonstrating that fusing their features creates a superior backbone for spatial tasks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach two different types of robots how to understand a physical room. One robot is a Super Reader (a Vision-Language Model, or VLM), and the other is a Master Director (a Video Generation Model, or VGM).
The big question the paper asks is: Which robot is better at understanding "spatial intelligence"?
To answer this, the researchers didn't let the robots learn new tricks or take a final exam. Instead, they put the robots in a "frozen" state—like pausing a video game character right before they start playing—and asked a simple question: "What information is already stored in your brain right now?"
They tested these frozen brains on three specific tasks, using a tiny, lightweight "probe" (think of it as a quick quiz) to see what each robot could recall.
The Three Tests
The "What's That?" Test (Semantic Tagging):
- The Task: Look at a video and list the objects you see (e.g., "sofa," "door," "table").
- The Result: The Super Reader (VLM) crushed this test. Because it was trained by reading books and looking at pictures with captions, it knows exactly what objects are called and what they look like. The Master Director (VGM) was okay, but it often missed key items (like forgetting the sofa was there).
The "Who Belongs to Whom?" Test (Instance Grouping):
- The Task: If you see a red chair in three different camera angles, can you tell that it's the same chair in all three shots?
- The Result: The Super Reader won again. It's great at saying, "That pixel group is a chair, and that pixel group over there is also that same chair." The Master Director struggled a bit, sometimes merging different objects together or failing to keep them distinct.
The "Where is Everything?" Test (3D Geometry):
- The Task: Can you build a 3D map of the room? How deep is the shelf? How far away is the camera?
- The Result: The Master Director (VGM) took the gold medal here. Because it was trained to create videos from scratch, it learned that objects have to stay in the right place and move realistically. This gave it a super-strong sense of depth, distance, and 3D structure. The Super Reader knew what the shelf was, but its 3D map was blurry and fuzzy.
The Big Discovery: The Perfect Team
The most exciting finding is that neither robot is perfect on its own. They are like two halves of a whole:
- The Super Reader is the expert on names and identity (Semantics).
- The Master Director is the expert on shape and space (Geometry).
The researchers tried a "naive fusion" (a simple mix-and-match). They took the frozen brain of the Super Reader and the frozen brain of the Master Director and glued them together.
The Result? The combined robot was a superhero. It could name every object perfectly and build a perfect 3D map of the room. The paper suggests that the best way to build future AI for robots or self-driving cars isn't to pick one model, but to combine the "brain" of a language expert with the "brain" of a video creator.
Summary in a Nutshell
- VLMs (Super Readers): Great at knowing what things are.
- VGMs (Master Directors): Great at knowing where things are in 3D space.
- The Solution: Mix them together to get an AI that understands both the names and the layout of the world perfectly.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.