Self Gradient Forcing: Native Long Video Extrapolation
This paper proposes Self Gradient Forcing (SGF), a two-pass training strategy that bridges the historical context-gradient gap in autoregressive video diffusion by enabling future losses to supervise how earlier generated latents are encoded into causal memory, thereby significantly improving native long-video extrapolation in subject identity, consistency, and temporal stability.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to paint a movie scene, one frame at a time. The robot is an artist that looks at what it just painted, adds a little bit of new detail, and then paints the next moment. This is how modern AI video generators work: they build long stories by chaining together tiny, self-generated steps. But there's a tricky problem. When the robot is learning, it usually practices by looking at perfect human-made videos (like a teacher showing the right answer). However, when it actually makes a movie, it has to rely on its own imperfect drawings from the previous step. This mismatch is like practicing piano with a metronome but playing a concert without one; the robot gets confused and starts to hallucinate, forgetting who the characters are or where the background was.
To fix this, scientists developed a method called "Self Forcing," where the robot practices by looking at its own messy drawings instead of perfect ones. This helps it get used to its own mistakes. But there's a hidden flaw in this practice: while the robot learns how to read its old drawings to make new ones, it doesn't learn how to write those old drawings into its memory in the first place. It's like a student who learns how to read a history book but never learns how to write the notes that go into the book. As the movie gets longer, the notes get messy, and the story falls apart. The robot forgets the main character's face, the background changes randomly, and the camera jumps around.
This is where a new idea called Self Gradient Forcing (SGF) comes in. The researchers behind this paper realized that to make long videos truly consistent, the robot needs to learn not just how to read its past, but how to write its past into a useful memory format. They proposed a clever two-step training trick. First, the robot generates a video normally, just like it would in real life, but without saving any "learning notes" (gradients) from that process. Then, in a second step, the robot takes a snapshot of that messy video and re-processes it in a special way that does allow it to learn how to write better notes for the future. It's like the robot makes a rough draft, throws it away, and then immediately writes a "study guide" based on that draft to teach itself how to remember things better next time.
The results are impressive. Using this method, the AI can take a video trained on just 5 seconds of footage and successfully extend it to 60 seconds or even 240 seconds (4 minutes) without the characters turning into monsters or the scenery dissolving into chaos. In tests, the videos made with SGF kept the subject's identity, the background layout, and the camera movement much more stable than the previous methods. The researchers found that while the old method eventually got confused and drifted apart, SGF kept the story coherent, proving that teaching the AI how to "write its own history" is the secret to making long, believable movies.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.