Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation
Proprio is a training-free framework that enhances the physical plausibility of frozen video generators by leveraging their internal flow residuals under latent perturbations as a self-scoring signal for inference-time refinement and selection.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a magical artist who can paint moving pictures (videos) from scratch. This artist is incredibly talented at making things look beautiful and realistic. However, they have a blind spot: they don't always understand how the real world works. They might paint a ball that floats upward instead of falling down, or a cup that shatters before it hits the floor.
For a long time, to fix these mistakes, people tried to hire a "physics teacher" (an external AI or a human) to watch the artist and say, "No, that's wrong." But the paper argues that this is like asking a different person to grade the artist's homework. The teacher might have their own biases, or they might not speak the same "language" as the artist, leading to misunderstandings.
Enter "Proprio."
The authors propose a new way to help the artist fix their own mistakes without hiring a teacher. They call this method Proprio, named after the biological sense of proprioception—the ability you have to know where your arm is without looking at it. Just as you can feel your own movement, Proprio lets the video generator "feel" its own creation to see if it makes sense.
Here is how it works, using simple analogies:
1. The "Shake Test" (Self-Scoring)
Imagine the artist has just finished a painting. Instead of asking a teacher to look at it, the artist gives the painting a gentle, controlled shake.
- The Logic: If the painting is physically correct (like a ball falling), a small shake won't confuse the artist. The artist's internal rules about how things move will still hold up.
- The Glitch: If the painting is wrong (like a ball floating), that small shake will cause the artist's internal logic to stumble. The artist will get confused about how to "fix" the image because the rules of physics are broken.
- The Score: Proprio measures this confusion. It calculates a "score" based on how much the artist stumbles when the video is slightly perturbed. A low score means the video is stable and physically plausible. A high score means the video is shaky and likely wrong.
2. The "Spotlight" (Dynamic Masking)
Not every part of a video needs to be checked equally. If a video shows a person sitting still, there's nothing to check. But if a car is crashing, that's where the physics matter.
Proprio uses a dynamic spotlight (a mask) to focus only on the moving parts of the video. It ignores the static background and zooms in on the action, asking, "Is this specific movement making sense?" This makes the "Shake Test" much more accurate.
3. Two Ways to Fix the Art
Once the artist has a score, Proprio offers two ways to improve the video:
The "Best of N" Search (Picking the Winner):
Imagine the artist paints 16 different versions of the same scene. Proprio acts like a quick judge, giving each version a "Shake Test" score. It then picks the one with the lowest score (the least confused) and says, "This is the winner." It's like rolling a die 16 times and picking the best roll.The "Self-Refinement" (Polishing the Art):
Instead of just picking a winner, Proprio can actually tweak the video. It goes back to the very beginning—the "noise" or the raw ingredients the artist used to start the painting. It slightly adjusts these ingredients and asks the artist to repaint the scene, hoping to lower the confusion score. It's like a sculptor chipping away a tiny bit of stone to make the statue stand more naturally, without changing the sculptor's tools or training.
Why This Matters
The paper shows that this "self-checking" method works better than asking external AI teachers (like large language models) to judge the videos.
- It's Native: The artist is judging its own work using its own internal rules, so there's no language barrier.
- It's Training-Free: You don't need to retrain the artist or teach them new physics. You just use the artist's existing brain to find and fix errors.
- The Results: When tested on videos of balls rolling, objects colliding, and liquids pouring, Proprio successfully picked and created videos that humans rated as more physically realistic.
In summary: Proprio gives the video generator a "gut feeling" about its own work. By gently shaking its creations and measuring the confusion, it can either pick the best version or subtly adjust the ingredients to make the physics feel right, all without needing an outside teacher.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.