← Latest papers
🤖 AI

Motion Attribution for Video Generation

The paper introduces Motive, a scalable, motion-centric data attribution framework that isolates temporal dynamics to identify high-influence training clips, thereby enabling data curation that significantly improves motion smoothness and physical plausibility in video generation models.

Original authors: Xindi Wu, Despoina Paschalidou, Jun Gao, Antonio Torralba, Laura Leal-Taixé, Olga Russakovsky, Sanja Fidler, Jonathan Lorraine

Published 2026-07-07
📖 5 min read🧠 Deep dive

Original authors: Xindi Wu, Despoina Paschalidou, Jun Gao, Antonio Torralba, Laura Leal-Taixé, Olga Russakovsky, Sanja Fidler, Jonathan Lorraine

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: The "Black Box" of Moving Pictures

Imagine you have a super-smart robot chef that can cook any dish you ask for. If you ask for a "spicy curry," it makes one. If you ask for "sushi," it makes that too. But what if you ask for a "dancing banana"? The robot might make a banana that looks like a banana but just slides across the plate instead of dancing.

In the world of AI video generation, we have these "robot chefs" (models) that can create videos. They are getting better at making things look real. However, scientists didn't really know which specific clips of video the robot was "tasting" to learn how to make things move.

  • The Issue: If the robot learns to make a ball bounce, is it because it watched a video of a basketball? Or a video of a rubber ball? Or maybe a video of a bouncing house?
  • The Old Way: Previous methods tried to figure out which training videos mattered, but they mostly looked at what things looked like (static appearance). They couldn't tell the difference between a video of a still painting and a video of a person dancing, even if both had a "person" in them. They treated motion like just another color on the canvas.

The Solution: Motive (The Motion Detective)

The authors created a new tool called Motive (MOTIon attribution for Video gEneration). Think of Motive as a specialized detective that only cares about movement, not the scenery.

Here is how Motive works, step-by-step:

1. Ignoring the Background (The "Motion Mask")

Imagine you are watching a movie where a car drives past a beautiful mountain.

  • Old Detective: "I see a car! I see a mountain! Both are important!"
  • Motive: "I don't care about the mountain. It's just sitting there. I only care about the wheels spinning and the car moving."

Motive uses a special "mask" (like a highlighter) that ignores static parts of the video (like the sky or a wall) and only highlights the parts that are actually moving. This allows it to calculate exactly which training videos taught the AI how to move things, rather than just how to draw them.

2. The "Recipe Book" Audit

The AI model is trained on a massive library of millions of videos. Motive goes through this library and asks: "If I remove this specific video, will the AI forget how to make a ball bounce?"

It gives every video in the library a score:

  • High Score: This video is a "masterclass" in movement. It taught the AI how to make things float, roll, or explode.
  • Low Score: This video is boring or confusing. It might have a lot of motion, but it's just a shaky camera, or it's a cartoon that moves in a way that breaks physics.

3. The "Taste Test" (Fine-Tuning)

The researchers didn't just stop at finding the good videos; they tested them.

  • They took a standard AI model.
  • They gave it a tiny, carefully selected diet of only the top 10% of videos that Motive said were the best for teaching motion.
  • The Result: The AI became much better at making videos that moved smoothly and obeyed the laws of physics.
    • It beat the "random diet" (where the AI ate whatever videos it found).
    • It even beat the "full diet" (where the AI ate all the videos), but using only 10% of the data.

Why This Matters (The "Aha!" Moment)

The paper claims this is the first time anyone has successfully separated "motion" from "appearance" in video AI.

  • Before: If you wanted an AI to learn how water flows, you might accidentally feed it videos of blue paint (which looks like water but doesn't flow). The AI would learn the wrong thing.
  • Now: Motive can say, "No, that blue paint video is useless for teaching flow. But that video of a river? That's the one that taught the AI how to make water flow."

The Results in Plain English

The team tested Motive on a benchmark called VBench (a report card for video AI).

  • Motion Smoothness: The videos looked less jittery and more fluid.
  • Dynamic Degree: The videos felt more alive and energetic.
  • Human Vote: When real humans watched the videos made by the Motive-trained AI, 74% of the time they preferred it over the original, untrained AI. They preferred it over the "full diet" AI 53% of the time.

The Bottom Line

Think of training a video AI like training a dog.

  • The Old Way: You throw a million treats at the dog and hope it learns to sit. Some treats are good, some are bad, and you don't know which ones worked.
  • The Motive Way: You use a special tool to identify the exact treats that made the dog sit perfectly. Then, you feed the dog only those specific treats. The dog learns faster, better, and with less food.

Motive proves that quality is better than quantity. By finding the specific videos that teach the AI how to move, we can build better video generators without needing to process the entire internet.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →