ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing
This paper introduces ElasticTTT, a novel test-time tuning framework that resolves the issue of prior collapse in video editing by employing Target Distribution Regularization, Contrastive CFG, and an Asynchronous Noise Schedule to preserve the generative prior and achieve state-of-the-art one-shot editing performance.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart robot artist that has spent years watching millions of movies. It knows exactly how to draw a cat, a car, or a sunset because it has memorized the "vibe" of all those scenes. This is the world of diffusion models, a type of artificial intelligence that creates images and videos by starting with static noise and slowly "denoising" it until a clear picture emerges. Think of it like a sculptor chipping away at a block of marble; the AI chips away at random static until a beautiful video appears.
Now, imagine you want to change just one thing in a video this robot made—maybe turn a sunny day into a rainy one, or swap a dog for a cat. Usually, the robot is stubborn; it wants to keep doing what it was trained to do. To fix this, scientists use a trick called Test-Time Tuning (TTT). This is like giving the robot a quick, intense crash course right before it starts drawing, using your specific video as the textbook. The robot learns your video's style so well that it can edit it perfectly. But here's the catch: if you study too hard on just one textbook, you might forget everything else you ever learned. The robot gets so obsessed with your specific video that it stops listening to your new instructions. It gets stuck in a loop, just replaying your original video over and over, or it gets confused and mixes up different parts of the scene. This is a problem scientists call "Prior Collapse."
This paper introduces a clever new framework called ElasticTTT to solve this exact problem. The researchers found that when you try to teach the AI to edit a video by fine-tuning it on that single video, the AI gets "rigid." It memorizes the video so tightly that it loses its flexibility, or "elasticity," to create something new. To fix this, they came up with three playful but powerful tricks. First, they added a little bit of "controlled chaos" to the training process. Instead of letting the robot memorize the video perfectly, they made the target slightly fuzzy, forcing the robot to stay flexible and not get stuck in a rigid memory loop. Second, they taught the robot to actively "push away" from the original video's style when it tries to follow your new instructions, ensuring it doesn't accidentally keep the old look. Finally, they gave the robot a special schedule where it treats the parts of the video you want to change differently from the parts you want to keep. It's like telling the robot, "Be wild and creative with the sky, but be very careful and steady with the ground."
The result is a system that can edit videos with incredible precision. The team tested ElasticTTT on 125 different video editing tasks, ranging from changing backgrounds to adding new objects. They found that their method consistently outperformed other top-tier editing tools. In fact, when they tested it on a larger, more powerful AI model, the improvements were even more dramatic, boosting the overall quality score from 4.91 to 7.00. The paper suggests that by keeping the AI "elastic" and preventing it from collapsing into a rigid memory of the original video, we can finally get these powerful robots to listen to our instructions without losing their creative spark. It's a step toward making video editing as easy as talking to a friend, without the robot getting confused or stubborn.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.