← Latest papers
💻 computer science

4Director: Controlling Video World Models with Rigid 3D Geometry

The paper introduces 4Director, a video world model that achieves precise camera and object control by conditioning generation on explicit 4D rigid geometry and a Motion Adapter, validated by a new dataset and evaluation metric to outperform existing methods in visual quality and motion consistency.

Original authors: Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

Published 2026-10-02
📖 4 min read☕ Coffee break read

Original authors: Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where a computer can watch a single photograph and then imagine a movie playing out from that still moment. This is the promise of video world models, a rapidly advancing field in artificial intelligence where machines learn to predict how a scene will change over time. For years, researchers have taught these systems to generate moving images by showing them thousands of hours of video, hoping the computer learns the rules of physics and motion on its own. However, a significant hurdle has remained: while these computers are becoming better at creating realistic movement, they are notoriously difficult to direct. If a user wants a specific object to turn around or a camera to swoop around a corner, the computer often ignores the instruction or creates a result that looks plausible but is geometrically wrong. The object might shrink instead of moving away, or it might blur into a ghostly smear as the camera moves. The core challenge has been bridging the gap between a human's clear, three-dimensional intent and the computer's tendency to guess based on flat, two-dimensional patterns.

A team of researchers from Stability AI and the University of Illinois Urbana-Champaign has introduced a new approach called 4Director to solve this problem. Instead of asking the computer to guess the path of an object, 4Director gives the computer a complete, three-dimensional map of the scene before the video generation even begins. In this system, every object in a starting photograph is reconstructed as a solid, digital 3D model, much like a sculptor creating a clay figure that has a front, a back, and sides, even if the original photo only showed the front. The background is also lifted into a cloud of points in space. Once this digital world is built, the user can draw a path for the camera and a path for the object, just as a film director would plan a shot. The computer then moves these solid 3D models along the prescribed paths with perfect rigidity, ensuring that if an object turns, it turns as a whole unit, and if the camera moves, the background shifts correctly.

The innovation lies in how this rigid skeleton is used to guide the final video. The system first renders a simple, black-and-white depth map of the scene, showing only the distance of every surface from the camera as the objects move along their paths. This depth map acts as a strict scaffold, defining exactly where every pixel should be in terms of space and time. A specialized training module, which the researchers call a Motion Adapter, then takes this rigid skeleton and fills in the missing details. It learns to paint the surface textures, the lighting, and the subtle, non-rigid movements—like the sway of a person's arm or the flutter of a flag—on top of the fixed 3D structure. This ensures that the final video follows the user's exact directions for position and orientation while still looking like a natural, living scene.

To teach this system how to work, the researchers had to create a massive new dataset because no existing collection of videos came with the necessary 3D maps. They built an automated pipeline that took over 20,000 video clips and reverse-engineered them, turning each one into a rigid 3D scene with a camera path and object paths. This dataset, named RealCOD-Rigid, allowed the Motion Adapter to learn the difference between the rigid movement of a solid object and the flexible movement of real-world details. The results show a marked improvement over previous methods. In tests, 4Director was able to keep objects consistent and recognizable even when they turned completely around or moved out of the camera's view and came back, a task where other systems often failed by losing the object's identity or blurring it into the background.

The researchers also developed a new way to measure success, recognizing that simply checking if an object is in the right spot is not enough if the object itself has changed into something else. Their new metric checks both the position and the identity of the object, ensuring that a car turning a corner still looks like the same car. When compared to other leading video generation tools, 4Director consistently produced videos with higher visual quality and much more accurate control over both the camera and the objects within the scene. The system demonstrates that by giving the computer a clear, three-dimensional understanding of the world before it starts generating the video, it is possible to achieve a level of control that was previously out of reach, allowing users to direct the flow of a scene with the precision of a professional filmmaker.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →