Distilling Physical Priors into Streaming World Models
The paper introduces PhyS, a three-stage framework that distills physical priors from a curated dataset of real-world interaction videos into a lightweight streaming world model via physics-aware fine-tuning and temporal credit routing, significantly improving the physical coherence of generated video rollouts compared to existing methods.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to predict the future, but instead of looking at a crystal ball, you are showing it a movie. This is the world of "video generation," where artificial intelligence tries to create new frames of a video that look and move just like the real thing. For a long time, these AI models were like great artists who could paint beautiful pictures but had no idea how gravity, water, or bouncing balls actually work. They could make a ball look pretty, but if they tried to predict what happens when it hits the ground, they might make it float away or pass right through the floor. Scientists call these predictions "world models," and the goal is to make them "streaming," meaning the AI can keep generating the story frame-by-frame, forever, without needing to restart the whole movie every time. The big question is: How do we teach a robot to understand the invisible rules of physics so its future predictions don't break the laws of nature?
The paper you are about to read introduces a clever new training camp called PhyS (Physical Streaming) to solve exactly this problem. The researchers found that the usual way of teaching these AI models was flawed. It was like trying to teach a student to drive a car by first showing them a movie of a car driving in reverse, and then asking them to drive forward. The student (the AI) learned to recognize the car, but it forgot the rules of the road, like how to stop or turn. The authors argue that the old method makes the AI forget the "physics priors"—the basic, common-sense rules of how the world moves—because the training process wipes them out.
To fix this, the team built a massive library of 120,804 real-world videos showing things actually happening: water splashing, balls bouncing, ice melting, and soft squishy things getting squashed. They didn't just save the videos; they used a super-smart AI assistant to write a detailed "physics story" for every clip, describing exactly why things moved the way they did. They then used this library to teach a giant, slow-thinking AI (a 14-billion-parameter model) to become a "Physics Teacher." This teacher learned the deep rules of the universe.
Next, they played a game of "telephone" to teach a much smaller, faster AI (a 1.3-billion-parameter model) how to be a "Streaming Driver." The small AI had to learn to predict the future frame-by-frame, just like a real-time video game, by copying the Physics Teacher. But there was a catch: sometimes the small AI would still make mistakes, like making a ball bounce too high. To fix this, the team invented a new scoring system called Temporal Credit Routing (TCR). Think of this like a referee in a soccer game who doesn't just say "good game" at the end. Instead, the referee watches every single second of the match. If the player kicks the ball into the goal at minute 10, the referee gives credit specifically to the kick at minute 10, not to the whole game. This helps the AI learn exactly when and where it broke the laws of physics, so it can fix that specific moment next time.
The results are impressive. When they tested their new AI on a "Physics IQ" test, it scored 18.2% higher than the giant teacher model it was based on. Even more surprisingly, it beat other top methods by huge margins: 23.7% better at one type of test, 14.8% at another, and a massive 31.4% at a third. The paper suggests that by giving the AI a library of real-world physics stories and a referee that checks its work second-by-second, we can finally build video generators that don't just look real, but actually act real.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.