Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing
This paper introduces Goku, a large-scale dataset of 2 million instruction-aligned video editing pairs covering complex multi-task and structural manipulations, along with the Goku-Edit model and Goku-Bench benchmark, to advance instruction-based video editing capabilities beyond simple appearance changes.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a magic video editor that can change anything in a movie just by you typing a sentence like, "Make the dog run faster and turn the sky purple." For a long time, these magic tools were like clumsy apprentices: they could swap a shirt color or remove a stray object, but if you asked them to move a character across the room or change the camera angle, they would get confused, freeze, or make a mess.
The paper introduces Goku, a massive new project designed to train these magic editors to become true masters. Here is the breakdown of what they did, explained simply:
1. The Problem: The "One-Trick Pony" Datasets
Before this, the training data for these AI video editors was like a library full of books about only one topic: changing the color of a car. If you wanted to learn how to drive a car (move a subject) or change the scenery (camera movement), the library had no books. The AI models were stuck because they only knew how to do simple "paint-by-numbers" tasks.
2. The Solution: The "Goku" Library
The researchers built Goku, a massive dataset containing 2 million video pairs. Think of this as building a giant, high-tech training gym for AI.
- What's inside? Unlike previous libraries, this one has complex challenges. It includes videos where the camera pans around, where a person moves from one side of the room to another, and where multiple changes happen at once (e.g., "Swap the shirt AND add a hat").
- How did they make it? They didn't just film these videos; they built a robot factory to create them. They used advanced AI to break down complex requests into small, manageable steps. For example, to make a dog run, they first taught the AI how to make a dog run in a single image, then how to make that image move smoothly into a video.
- The Quality Control: To ensure the videos weren't garbage, they installed a "triple-check" security system. Every single video was checked three times by a super-smart AI (Gemini) to make sure the instructions were followed, the movement looked real, and the video didn't glitch. They threw away about 88% of the attempts to keep only the best 2 million.
3. The New Tool: Goku-Edit
With this new library, they built a new video editor called Goku-Edit.
- The Brain: Instead of just reading text, this editor uses a "Multimodal Large Language Model" (MLLM). Think of this as giving the editor a brain that can actually understand the story and the instructions, not just match keywords.
- The Two-Handed Approach: The biggest innovation is that the editor has two "hands" working together:
- Hand A (The Architect): This hand draws a rough map (a mask) of where the changes should happen. It focuses on structure and movement.
- Hand B (The Painter): This hand focuses on the details, colors, and textures.
- Why this matters: By separating the "where" from the "what," the editor doesn't get confused. It knows exactly where to move the dog without accidentally painting the dog's fur blue.
- The "Spatial CFG": This is a fancy term for a safety net. It ensures that when the editor makes a change, it doesn't accidentally spill over and ruin the rest of the video. It keeps the changes tight and precise.
4. The Test: Goku-Bench
To prove their new editor is the best, they created a new test called Goku-Bench.
- Instead of simple tests, they gave the AI 1,000 difficult challenges, like "Move the camera to the left while the cat jumps."
- They used 7 new ways to grade the results, checking things like "Did the physics make sense?" and "Did the camera move smoothly?"
- The Result: Goku-Edit beat all other open-source models, improving its ability to follow instructions by up to 8%. It was particularly good at moving things around and handling complex, multi-step requests where other models failed.
Summary
In short, the paper says: "We built a massive, high-quality training library (Goku) that teaches AI how to do complex video editing, not just simple color swaps. We built a new editor (Goku-Edit) that uses a two-part system to understand structure and details separately, and we proved it works better than anything else on a new, harder test (Goku-Bench)."
They did not claim this technology is ready for hospitals or specific clinical uses; they focused entirely on making video editing smarter, more flexible, and more capable of handling complex creative requests.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.