← Latest papers
💻 computer science

ProPhy: Progressive Physical Alignment for Dynamic World Simulation

ProPhy is a progressive physical alignment framework that enhances the physical consistency of dynamic world simulations by employing a two-stage Mixture-of-Physics-Experts mechanism to extract fine-grained physical priors and transfer vision-language reasoning capabilities for anisotropic, physics-aware video generation.

Original authors: Zijun Wang, Panwen Hu, Jing Wang, Terry Jingchen Zhang, Yuhao Cheng, Long Chen, Yiqiang Yan, Zutao Jiang, Hanhui Li, Xiaodan Liang

Published 2026-04-15
📖 5 min read🧠 Deep dive

Original authors: Zijun Wang, Panwen Hu, Jing Wang, Terry Jingchen Zhang, Yuhao Cheng, Long Chen, Yiqiang Yan, Zutao Jiang, Hanhui Li, Xiaodan Liang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are asking a talented but slightly naive artist to paint a scene based on your description: "A basketball hits the ground, kicking up dust, while a cup of coffee sits nearby, steaming in the cold air."

The Problem with Current AI:
Most current video-generating AIs are like that artist who is great at painting colors and shapes but doesn't quite understand how the real world works.

  • If you ask for the basketball, they might paint the dust rising before the ball hits the ground.
  • If you ask for the coffee, they might make the liquid level rise magically because of the fire, or make the coffee pot catch fire instantly.
  • They treat the whole video like a single, blurry blob. They don't realize that the "dust rule" only applies to the basketball, and the "liquid rule" only applies to the coffee. They try to apply one set of "physics rules" to the entire picture at once, leading to chaos.

The Solution: ProPhy (The "Progressive Physical Alignment" Framework)
The researchers behind ProPhy decided to fix this by giving the AI a team of specialized experts and a strict training regimen. Think of ProPhy not as a single artist, but as a high-end film production studio with a very specific workflow.

Here is how ProPhy works, broken down into simple steps:

1. The "Script Readers" (Semantic Experts)

First, before the camera even starts rolling, ProPhy has a team of "Script Readers."

  • What they do: They read your text prompt ("basketball," "coffee," "snow") and immediately figure out the big picture physics.
  • The Analogy: Imagine a director looking at a script and saying, "Okay, this scene needs a Gravity Expert for the falling ball and a Fluid Dynamics Expert for the coffee. Let's get those specialists ready."
  • The Innovation: Unlike older models that just guess the physics, ProPhy explicitly asks, "What kind of physics are we dealing with here?" and assigns the right "expert" to the job.

2. The "Special Effects Crew" (Refinement Experts)

Once the big picture is set, the real magic happens. ProPhy doesn't just apply the "Gravity Expert" to the whole screen. It breaks the video down into tiny pieces (like pixels or small tiles) and assigns a specific expert to each piece.

  • What they do: They look at the basketball and say, "You need to bounce and kick up dust." They look at the coffee and say, "You need to stay liquid and not catch fire."
  • The Analogy: Imagine a massive construction site. Instead of one foreman telling everyone to "build a wall," you have a Carpenter working on the wood, a Plumber working on the pipes, and an Electrician on the wires. They all work on their specific parts of the building simultaneously, ensuring the wood doesn't turn into water and the pipes don't catch fire.
  • The Result: The dust only rises where the ball hits. The coffee stays safe. The physics are "fine-grained" and precise.

3. The "Reality Check" (VLM Alignment)

How do we teach the AI to be this precise? The researchers used a clever trick involving Vision-Language Models (VLMs).

  • The Problem: The video generator is bad at pointing out exactly where a physical event happens.
  • The Solution: They used a different type of AI (a VLM) that is really good at looking at a video and describing it. They asked the VLM: "Where exactly is the dust rising?" and "Where is the coffee?"
  • The Analogy: Think of the VLM as a strict physics teacher. The video generator is the student. The teacher points at the video and says, "No, the dust is here, not there. The coffee is here, not there." The student (ProPhy) then copies the teacher's notes to learn exactly where to apply the rules.

Why This Matters (The "World Simulator")

Current AI is like a dream. In a dream, you can fly, water can turn into fire, and gravity can be optional. It looks pretty, but it's not real.

ProPhy is trying to turn that dream into a simulation.

  • Old AI: "Here is a video of a ball. It looks like a ball." (But it might float away).
  • ProPhy: "Here is a video of a ball. It hits the ground, obeys gravity, kicks up dust, and stops. It behaves exactly like a real ball would."

Summary

ProPhy is a new way to teach AI how to make videos that obey the laws of physics. Instead of guessing, it:

  1. Identifies the specific physics rules needed (Gravity, Fluids, Fire).
  2. Assigns a specialist to handle those rules for every tiny part of the video.
  3. Trains using a "teacher" AI that points out exactly where things happen, ensuring the dust rises only when the ball hits, and the coffee doesn't spontaneously combust.

It's the difference between a magic trick and a science experiment. ProPhy wants to build a World Simulator where the video generation is as reliable as reality itself.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →