← Latest papers
🤖 AI

PhysVid: Physics Aware Local Conditioning for Generative Video Models

PhysVid introduces a physics-aware local conditioning scheme that annotates temporally contiguous frame chunks with grounded descriptions and employs negative physics prompts to significantly enhance the physical plausibility of generative video models.

Original authors: Saurabh, Pathak, Elahe Arani, Mykola Pechenizkiy, Bahram Zonooz

Published 2026-03-30
📖 4 min read☕ Coffee break read

Original authors: Saurabh, Pathak, Elahe Arani, Mykola Pechenizkiy, Bahram Zonooz

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot artist how to paint a moving picture (a video) based on a simple sentence like, "A car drives down a rainy street."

Current AI video generators are like incredibly talented but slightly clumsy artists. They can make the car look beautiful and the rain look shiny, but they often mess up the physics. The car might float like a ghost, the rain might fall upward, or the wheels might spin backward while the car moves forward. They know what things look like, but they don't really understand how things work.

PhysVid is a new method that teaches the AI to be a better "physics student" so it doesn't just guess; it understands the rules of the universe.

Here is how it works, broken down into simple concepts:

1. The Problem: The "Global" vs. The "Local"

Imagine you are directing a movie.

  • Old Way (Global Prompt): You give the director one big instruction for the whole movie: "Make a movie about a car driving in the rain."
    • The Issue: The director gets the general vibe, but when it comes to the specific moment the car hits a puddle, they might forget the rules of water. They might make the water splash upward or the car slide sideways like it's on ice, even if the road is dry. The instruction was too broad to catch the tiny, split-second details.
  • PhysVid's Way (Local Conditioning): Instead of one big instruction, PhysVid breaks the movie into tiny 1-second clips. For each clip, it gives the director a specific, physics-focused note.
    • Clip 1: "The car moves forward; the wheels spin forward."
    • Clip 2: "The car hits a puddle; water splashes outward and downward due to gravity."
    • Clip 3: "The rain falls straight down; the windshield wipers push water sideways."

By focusing on these tiny chunks, the AI pays attention to the specific rules that apply right now, rather than just the general theme.

2. The Secret Sauce: The "Physics Tutor" (VLM)

How does the AI know what those specific notes should be? It uses a Vision Language Model (VLM), which acts like a Physics Tutor.

  • During Training: The AI watches thousands of real videos. The Tutor looks at a 1-second clip and says, "Okay, in this specific second, the ball is falling because of gravity, and the shadow is getting longer because the sun is moving." It writes these rules down as a special "physics prompt."
  • The Result: The AI learns to associate the visual of a falling ball not just with the word "ball," but with the specific rule "gravity pulls it down."

3. The "Anti-Gravity" Trick (Counterfactuals)

This is the cleverest part. To teach the AI what not to do, PhysVid uses a "Devil's Advocate" approach.

  • The Positive Prompt: "The ball falls down."
  • The Counterfactual (Negative) Prompt: The AI is also asked to imagine the wrong version: "The ball floats up," or "The ball turns into a square while falling."

During the generation process, the AI is told: "Make the video look like the Positive Prompt, but make sure it looks NOTHING like the Negative Prompt."

Think of it like a teacher correcting a student: "Don't draw the car floating! Draw it on the road!" By explicitly showing the AI what "bad physics" looks like, it learns to steer clear of those mistakes.

4. The Result: A Smarter, Smaller Artist

The paper shows that PhysVid is a 1.7 billion parameter model. For context, the giant model it is compared to (Wan-14B) is 14 billion parameters—almost 10 times bigger.

  • The Analogy: Imagine a giant encyclopedia (the big model) that knows everything but is slow and sometimes gets confused by the details. PhysVid is like a small, sharp-witted detective who only cares about the rules of physics.
  • The Outcome: Even though PhysVid is much smaller, it generates videos that obey the laws of physics 33% better than the massive models. The water flows correctly, the cars drive realistically, and objects don't magically disappear or change shape.

Summary

PhysVid is like giving a video generator a magnifying glass and a rulebook. Instead of looking at the whole movie at once, it zooms in on every second, checks the physics rules for that specific moment, and uses a "what-if" game to avoid silly mistakes. The result is a video that feels real because the AI finally understands that gravity pulls things down, not up.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →