← Latest papers
💻 computer science

When Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video Generation

This paper introduces Physical and Semantic Direct Preference Optimization (PSDPO), a novel method that resolves the trade-off between physical plausibility and semantic consistency in text-to-video generation by modulating preference pair contributions based on signal agreement, thereby improving physical realism without sacrificing semantic fidelity.

Original authors: Siwei Meng, Yawei Luo, Shu Zhang, Ping Liu

Published 2026-07-21
📖 4 min read☕ Coffee break read

Original authors: Siwei Meng, Yawei Luo, Shu Zhang, Ping Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to paint pictures based on your spoken descriptions. You tell it, "Draw a cat chasing a laser pointer," and it spits out a masterpiece. But then you ask for something slightly more complex, like "A cat jumping off a table." The robot might draw a cat that defies gravity, floating in mid-air like a ghost, or one that melts into the floor. This is the current state of Text-to-Video generation. These AI models are incredible at making things look real and matching the words you type, but they often struggle with the invisible rules of our universe: physics. They don't always understand that objects have weight, that liquids splash, or that things can't pass through each other.

To fix this, researchers have been using a technique called Direct Preference Optimization (DPO). Think of this like a strict art teacher. You show the teacher two videos: one where a ball bounces realistically, and one where it floats. The teacher says, "I like the bouncing one better." The AI learns from this feedback to make more bouncing balls. However, there's a catch. Sometimes, the "better" video for physics is actually a disaster for the story. The bouncing ball video might be physically perfect, but it might be a red ball when you asked for a blue one, or it might be a dog instead of a cat. The AI, trying to please the teacher's physics lesson, starts ignoring your original request entirely. It becomes a master of physics but a terrible storyteller.

This is the puzzle tackled by a new paper from researchers at the University of Nevada, Reno, and Zhejiang University. They noticed that when AI tries to learn physics from these "this is better than that" comparisons, it often gets confused and starts hallucinating, dropping the semantic meaning (the story) to focus solely on the movement. They call this the "physical-semantic conflict."

The researchers propose a clever solution called PSDPO (Physical and Semantic Direct Preference Optimization). Instead of blindly listening to every piece of feedback, PSDPO acts like a smart filter. It checks the feedback before letting the AI learn from it. If the AI is shown a video that is physically perfect but semantically wrong (like the red ball instead of the blue one), PSDPO says, "Hold on, this lesson is too confusing; let's not learn from this one right now." It weighs the lessons, giving full credit to videos that are both physically correct and match the story, while gently downgrading the ones that only get the physics right but mess up the story.

To make this work, they built a special "judge" system called the Physical Commonsense Evaluation (PCE). This system doesn't just look at the video; it breaks the scene down into steps, checking if the physics makes sense (like does the water flow down?) and if the story matches (is it actually a cup of water?). They found that this relative comparison (Video A vs. Video B) is much more reliable than trying to grade a single video on its own.

The results are promising. When they tested their new method, the AI generated videos that were up to 2 times better at following the laws of physics compared to the standard methods, all while keeping the story exactly as the user asked. They also discovered that the order in which the AI learns matters. If you teach it the hard, confusing lessons too early, it gets lost. But if you start with the clear, easy lessons (where physics and story agree) and only introduce the tricky, conflicting ones later, the AI learns much faster and stays on track.

In short, the paper suggests that by being a bit more selective about which lessons the AI learns from—and by teaching in the right order—we can create video generators that are not only physically realistic but also faithful to the stories we tell them. It's a step toward AI that understands not just how things move, but what things are, ensuring the magic of the story isn't lost in the pursuit of realism.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →