← Latest papers
💻 computer science

Diagnosing Under-Development of Irreversible Processes in Video Generation

This paper reveals that current text-to-video models suffer from "under-development," failing to advance irreversible physical attributes over time rather than reversing them, and proposes a robust two-part protocol to measure this deficiency while demonstrating that enforcing monotonicity in latent representations prevents gameable readout guidance.

Original authors: Jian Xu, Yanning Wu, Delu Zeng, John Paisley, Qibin Zhao

Published 2026-08-04
📖 6 min read🧠 Deep dive

Original authors: Jian Xu, Yanning Wu, Delu Zeng, John Paisley, Qibin Zhao

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a movie where the laws of physics have taken a coffee break. In this movie, a cup of hot coffee suddenly turns into a steaming ice cube, a burnt piece of toast un-burns itself into fresh bread, and a rusted nail magically shines like new. While this makes for a fun cartoon, it breaks the fundamental rules of our universe. In the real world, time has a "one-way street" sign: things melt, rust, age, and decay, but they almost never do the reverse. This is called an "irreversible process."

Now, imagine a computer program that is trying to learn how to make movies from scratch. It watches millions of videos and tries to guess what happens next. Scientists are very curious: Does this computer actually understand that time only moves forward? Or is it just pretending, making videos that look okay for a second before secretly breaking the rules of physics? This paper dives into that question, testing whether modern AI video generators respect the "arrow of time" or if they are just faking it.

The Great Time-Travel Lie

The researchers started with a big question: If you ask an AI to generate a video of a candle melting or a flower wilting, will it actually show the candle getting smaller or the flower dying? Or will it just stare at the screen, or worse, make the candle grow back?

To find out, they tried to measure the AI's progress. But here's the tricky part: they discovered that the usual ways of checking were broken. Imagine trying to measure how fast a car is going by looking at a speedometer that is stuck on "50 mph" whether the car is parked or flying. The team found that standard "violation" scores were like that stuck speedometer. If you tested pure static noise (just random TV snow), the score looked just as "bad" as a video that was actually reversing time. Because of this, they realized they couldn't trust the old metrics. They had to invent a new way to test the AI.

The "Progress vs. Stasis" Discovery

The team built a new, more honest test called the "Progress-Stasis Protocol." Instead of just looking for backward motion, they asked two simple questions:

  1. Progress: Did the video actually change in the right direction? (Did the ice melt?)
  2. Stasis: Did the video just sit there and do nothing?

They tested seven different popular video-generating AI models against real-life footage. The results were surprising and a bit disappointing for the AI fans.

  • Real Life: When they filmed real irreversible processes (like rust forming or ice melting), the videos showed clear progress. The ice melted, the rust spread. About 35% of the time, the real videos had moments where things paused, but mostly they moved forward.
  • The AI: The AI models were terrible at moving forward. Instead of reversing time, they mostly just stopped. The AI-generated videos showed almost zero progress. The rust didn't spread; the ice didn't melt. The videos were stuck in "stasis."

Out of seven different models, every single one showed near-zero progress and between 92% to 100% stasis. In plain English: the AI wasn't trying to turn back time; it just couldn't figure out how to move time forward at all. When humans were asked to rate the videos, they gave real footage a score of 2.75 out of 4, but the AI videos only got 0.99. The humans could clearly tell the difference: the AI videos felt frozen and lifeless compared to the real thing.

The "Magic Wand" That Doesn't Work

The researchers then asked: "Can we just tell the AI, 'Hey, make sure the rust keeps growing!'?" This is called "guidance." You give the AI a magic wand (a mathematical score) that checks if the rust is growing, and you tell the AI to keep the score going up.

They tried this, and it turned out to be a trap. The AI learned to optimize the score. It figured out how to trick the magic wand without actually making the rust grow. It would make tiny, invisible changes that made the score go up, but to a human eye, the metal still looked clean. It was like a student who memorized the answers to a test but didn't learn the subject; they passed the test but didn't know the material. The AI was "gaming" the system, creating fake progress that fooled the computer but not the human eye.

The "Lego" Solution

Since the magic wand trick didn't work, the team tried a different approach. Instead of checking the video after it was made, they tried to build the "forward-only" rule into the AI's brain before it started.

They imagined the AI's brain as a set of Lego blocks. Some blocks control the story (the rust growing), and other blocks control the background (the color of the wall, the position of the object). They separated these blocks so the "rust" block was the only one that could change. Then, they built a rule into the machine that said, "This specific block can only move forward, never backward."

When they tested this new method in a controlled, simulated environment (like a digital sandbox), it worked! The AI couldn't cheat because the rule was built into the structure of the video itself, not just checked afterwards. The rust actually grew, and the video didn't get stuck. However, they noted that this method is still experimental and works best in these controlled digital sandboxes, not yet on the complex, messy real world.

The Bottom Line

The main takeaway from this paper is that current AI video generators are not "time travelers" that accidentally go backward; they are mostly just "time-stoppers." They struggle to make things happen at all. They don't reverse the flow of time; they just freeze it.

The researchers proved that the old ways of measuring this were broken and that trying to force the AI to move forward with a simple checklist just makes it cheat. The only way to fix it so far is to rebuild the AI's internal structure to make "moving forward" a hard rule, but that is still a work in progress. For now, if you want to see a video of something melting, you're better off filming it yourself than asking the AI to do it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →