InTraGen: Trajectory-controlled Video Generation for Object Interactions
InTraGen is a novel pipeline that enhances text-to-video generation of multi-object interactions by introducing a multi-modal encoding mechanism with object ID injection, supported by four new datasets and a trajectory quality metric to evaluate and improve visual fidelity and interaction realism.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you want to tell a story using a video generator. You type in a prompt like, "A red car hits a blue car, and then the blue car falls into a hole."
Current AI video generators are like talented but scatterbrained artists. They hear your story and try to paint it, but they often get the sequence wrong. They might paint both cars crashing immediately, or they might forget which car is which, turning the red car into a blue one halfway through. They struggle with the "when" and "who" of the story because text alone is too vague for complex timing.
InTraGen is the solution the authors propose. Think of it as giving the artist a precise choreography map instead of just a vague description.
Here is how it works, broken down into simple concepts:
1. The Problem: The "Blurry Script"
Text is great for ideas, but bad for physics. If you tell a human, "The ball bounces off the wall," they understand the timing. If you tell an AI, it might get confused about when the bounce happens or which ball is bouncing. It's like trying to conduct an orchestra by only humming the melody; the musicians (the AI) don't know exactly when to play their notes.
2. The Solution: The "Choreography Map" (Trajectories)
Instead of just text, InTraGen asks you to draw lines (trajectories) showing exactly where every object should move, frame by frame.
- The Input: You draw a line for the red car, a line for the blue car, and a line for the ball.
- The Magic: The AI doesn't just guess; it follows your map like a train on tracks.
3. The Secret Sauce: The "Name Tags" (Object IDs)
This is the paper's biggest innovation. Previous methods tried to follow the lines, but they got confused when two objects got close.
- The Old Way: Imagine two dancers moving close together. If the director only says "move left," the dancers might swap places or merge into a blob.
- The InTraGen Way: The system gives every object a permanent color-coded name tag (an Object ID). Even if the red car and blue car are right next to each other, the AI knows, "That's the Red Car's path, and that's the Blue Car's path." It keeps their identities separate, ensuring they don't morph into each other or disappear.
4. The Training Ground: The "Toy Box"
To teach this AI how to handle crashes, falls, and collisions, the authors didn't just use real videos (which are messy and hard to measure). They built a giant digital toy box using game engines like Blender and Unity.
- They created 50,000 videos of billiard balls colliding, dominoes falling, and football players kicking balls.
- Because these are computer-generated, they know exactly where every object is at every second. This is like having a perfect answer key to grade the AI's homework.
5. The Result: A Better Director
The paper shows that when you give the AI these "maps" and "name tags":
- It follows the rules: If you say the ball hits the wall at second 5, it hits the wall at second 5.
- It keeps identities: The red ball stays red; the blue ball stays blue.
- It looks real: The collisions and physics look much more natural than before.
The Bottom Line
InTraGen is like upgrading a video generator from a dreamer (who guesses what happens) to a precision engineer (who follows a blueprint). By combining a visual map of movement with a system to keep track of "who is who," it allows us to create videos of complex interactions—like a pool game or a car crash—that actually make sense and look realistic.
In short: It's the difference between telling a robot "make a mess" and handing it a blueprint for exactly how the mess should happen.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.