← Latest papers
💻 computer science

MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer

MoVieDrive introduces a unified diffusion transformer framework that leverages shared and specific components to simultaneously generate controllable, high-quality multi-modal (e.g., RGB, depth, semantic) and multi-view urban driving videos, addressing the limitations of existing single-modal or multi-model approaches in autonomous driving scene synthesis.

Original authors: Guile Wu, David Huang, Dongfeng Bai, Bingbing Liu

Published 2026-03-16
📖 4 min read☕ Coffee break read

Original authors: Guile Wu, David Huang, Dongfeng Bai, Bingbing Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a movie director for a self-driving car. Your job is to create millions of different driving scenarios to teach the car how to drive safely. You need to show it sunny days, snowy nights, rainy streets, and even weird "what-if" situations like a cow jumping in front of the car.

In the past, making these movies was like having a team of three separate artists:

  1. Artist A drew the picture (the RGB video).
  2. Artist B drew the depth map (a blueprint showing how far away things are).
  3. Artist C drew the semantic map (a coloring book showing what is a road, what is a car, and what is a tree).

The problem? These three artists didn't talk to each other. Sometimes Artist A drew a car, but Artist B forgot to draw it, or Artist C painted it the wrong color. Also, hiring three separate teams was expensive and messy.

Enter MoVieDrive. Think of MoVieDrive as a super-talented "Swiss Army Knife" director who can do all three jobs at once, perfectly in sync.

Here is how it works, broken down into simple concepts:

1. The "One-Brain" Approach (Unified Model)

Instead of hiring three different models, MoVieDrive uses one single brain (a Diffusion Transformer) to generate everything.

  • The Analogy: Imagine a chef who can cook a steak, bake a cake, and make soup simultaneously in one pot, rather than having three different chefs in three different kitchens. Because it's all one brain, the steak (the video), the cake (the depth), and the soup (the semantics) are perfectly matched. They never contradict each other.

2. The "Recipe Book" (Conditioning Inputs)

To tell this director what to film, you don't just say "make a movie." You give it a very specific recipe using three types of ingredients:

  • The Script (Text Prompts): You tell it, "It's a rainy Tuesday night in Tokyo."
  • The Blueprint (Layout Conditions): You give it a rough sketch of where the cars and roads should be (like a traffic map).
  • The Reference Photo (Context): You show it the first frame of the scene so it knows where to start.

MoVieDrive mixes all these ingredients together so the AI knows exactly what to build.

3. The "Shared & Specialized" Learning

This is the secret sauce. The model is built like a team of twins:

  • The Shared Twins (Modal-Shared): These parts learn the general rules of the world. They understand that cars move, rain falls, and time passes. They make sure the video looks smooth and logical over time.
  • The Specialized Twins (Modal-Specific): These parts are the experts. One twin focuses on making the colors look real (RGB), another on making the distance look real (Depth), and the third on making the object labels accurate (Semantic).

They work together in the same room. The "Shared" twins keep the story consistent, while the "Specialized" twins make sure the details are perfect for their specific job.

4. Why This Matters (The Magic Result)

Because this system generates everything at once, it creates perfectly synchronized data.

  • No Glitches: You won't see a car in the video that disappears in the depth map.
  • Long Videos: It can generate long, continuous driving movies (like 15 seconds or more) without needing a reference frame to hold its hand. It's like the director can improvise a whole scene without a script.
  • Weather & Time Travel: You can instantly change the weather from sunny to snowy just by changing the text prompt, and the AI will adjust the road, the cars, and the shadows instantly.

In a Nutshell

MoVieDrive is like upgrading from a chaotic film set with three disconnected crews to a single, magical AI director. This director can instantly generate a driving movie, a 3D depth blueprint, and a semantic map of the scene all at the same time, ensuring that every single detail matches perfectly. This helps self-driving cars learn faster and safer because they can practice in a perfect, simulated world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →