UniVideo: Unified Understanding, Generation, and Editing for Videos
UniVideo is a versatile dual-stream framework that unifies video understanding, generation, and editing under a single multimodal instruction paradigm, achieving state-of-the-art performance and enabling novel capabilities like task composition and zero-shot transfer from image editing data.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart assistant who can do three things at once: watch a video and tell you what's happening, create a brand new video from scratch, and edit an existing video just by listening to your instructions.
Until now, most AI assistants were like specialized workers. One was great at watching videos but couldn't make them. Another was a great video director but couldn't understand complex instructions. A third was a master editor but needed specific tools for every single job.
UniVideo is like hiring a "Swiss Army Knife" assistant who does all of these jobs in one package. Here is how it works, using simple analogies:
1. The Two-Brain System
The paper explains that UniVideo uses a "dual-stream" design, which is like having two specialized brains working together:
- The "Understanding Brain" (MLLM): Think of this as a very well-read librarian. It reads your instructions, looks at your reference photos, and watches your videos. It understands the story, the context, and the logic. It knows what a "sunny beach" looks like and what a "detective hat" implies.
- The "Creation Brain" (MMDiT): Think of this as a master painter or a special effects artist. It doesn't read words well, but it is incredible at taking raw materials and painting them into a moving picture.
The Magic Connection: The "Understanding Brain" reads your request and whispers a detailed plan to the "Creation Brain." The "Creation Brain" then paints the video, making sure the details (like the texture of a shirt or the movement of a car) look real and consistent.
2. What Can It Do?
Because these two brains are trained together, UniVideo can handle a huge variety of tasks without needing a different app for each one:
- The Director: You can say, "Make a video of a man in a Hawaiian shirt sitting on a beach," and it creates it.
- The Editor: You can upload a video and say, "Change the man's shirt to green," or "Replace the man with the person in this photo," and it does it.
- The Storyteller: You can show it a video and ask, "What is happening here?" and it will describe the scene in detail.
- The "Thinking" Mode: This is a special feature. Sometimes your instructions are messy or come with drawings. For example, you might draw a rough sketch on a photo and say, "Make this happen." The "Understanding Brain" figures out your messy drawing and turns it into a clear plan for the "Creation Brain" to follow.
3. The "Zero-Shot" Superpower
One of the coolest things the paper claims is that UniVideo can do things it was never explicitly taught to do.
Imagine you teach a child how to edit photos (changing a background from a park to a desert). Then, you show them a video of a park and ask them to change the background to a desert. Even if you never specifically taught them how to edit videos, they might figure it out because they understand the concept of changing a background.
UniVideo does this too. It was trained heavily on editing images, and when asked to edit videos (like changing the weather or the material of an object), it transfers that knowledge. It didn't need a specific "video editing" class for every single type of edit; it just applied what it learned from images to the video world.
4. Why Is This Different?
Previous models were like a toolbox where you had to pick the right tool for the job. If you wanted to edit a video, you needed a specific "video editor" tool. If you wanted to generate a video, you needed a "generator" tool.
UniVideo is like a universal remote control. You press one button (give one instruction), and the device figures out whether it needs to watch, create, or edit, and then does the job perfectly. The paper shows that this "universal" approach is actually better than using specialized tools for many tasks, especially when you want to combine tasks (like editing a video while changing the style of the characters).
In short: UniVideo is a single AI model that can watch, create, and edit videos by using a smart "reader" to understand your wishes and a talented "artist" to bring them to life, all while being able to figure out new tricks on its own.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.