← Latest papers
💻 computer science

YoCausal: How Far is Video Generation from World Model? A Causality Perspective

This paper introduces YoCausal, a novel benchmark leveraging temporally reversed real-world videos to evaluate video diffusion models, revealing that while they can perceive the arrow of time, they still lack genuine causal reasoning capabilities compared to human-level cognition.

Original authors: You-Zhe Xie, Yu-Hsuan Li, Jie-Ying Lee, Kaipeng Zhang, Yu-Lun Liu, Zhixiang Wang

Published 2026-05-29
📖 5 min read🧠 Deep dive

Original authors: You-Zhe Xie, Yu-Hsuan Li, Jie-Ying Lee, Kaipeng Zhang, Yu-Lun Liu, Zhixiang Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to understand how the world works. You show it a video of a glass shattering. A human instantly knows: the hammer hit the glass, then the glass broke. If you played that video backward, showing the shards flying up and reassembling into a whole glass, you would immediately think, "That's impossible!" Your brain is screaming, "Wait, that's wrong!"

This paper, YoCausal, asks a simple but profound question: Do today's advanced AI video generators actually "know" this is wrong, or are they just good at mimicking the look of a video without understanding the logic behind it?

Here is the breakdown of their discovery, using some everyday analogies.

The Problem: The "Magic Trick" vs. Real Understanding

Current AI video models are like incredibly talented magicians. They can make a video where a glass shatters look incredibly realistic. But if you ask them, "Did the glass break because of the hammer, or did the hammer appear because the glass broke?" they might not have a clue. They might just be memorizing patterns: "Usually, broken glass looks like this."

The researchers wanted to know if these AIs truly understand causality (cause and effect) or if they are just "statistical parrots" repeating what they've seen before.

The Solution: The "Backward Video" Test

To test this, the team created a new benchmark called YoCausal. They used a clever trick inspired by how we test babies: The Violation of Expectation.

  • The Baby Test: If you show a baby a video of a ball rolling normally, they watch calmly. If you play it backward (the ball rolling uphill on its own), the baby looks surprised. This surprise proves they understand how balls should move.
  • The AI Test: The researchers took real-world videos (like a person wiping a dirty plate) and played them backward. They didn't need to create fake videos; they just hit "reverse" on real footage. This is a "counterfactual" sample—a scenario that violates the laws of time and cause.

They then asked the AI: "Which version of this video feels more 'normal' to you?"

How They Measured "Surprise"

AI video models work by "denoising"—taking a blurry, static-filled image and cleaning it up step-by-step to make a clear video.

  • The Analogy: Imagine the AI is trying to solve a puzzle.
  • The Forward Video: The puzzle pieces fit together logically. The AI solves it easily (low "effort" or loss).
  • The Backward Video: The puzzle pieces are scrambled in a way that defies logic. The AI struggles to make sense of it (high "effort" or loss).

If the AI truly understands causality, it should struggle much more with the backward video than the forward one. This struggle is their metric for "surprise."

The Two-Level Test

The researchers realized that just being confused by backward videos isn't enough. An AI might just know that "time usually moves forward" without understanding why. So, they built a two-level test:

  1. Level 1: The "Time Arrow" Test (RSI)

    • Question: Does the AI know that time usually flows forward?
    • Result: Many advanced AIs passed this. They knew backward videos were weird. But this is like knowing that a movie should be watched from start to finish, not necessarily understanding the plot.
  2. Level 2: The "Story Logic" Test (CCI)

    • Question: Does the AI understand the story?
    • The Trick: They split videos into two groups:
      • Causal Videos: Things where A definitely causes B (e.g., a hammer hitting a nail).
      • Non-Causal Videos: Things where A doesn't really cause B, it just happens to be there (e.g., a car driving down a highway).
    • The Logic: If you reverse a causal video, it's a double violation (time is wrong + physics is wrong). If you reverse a non-causal video, it's only a single violation (time is wrong).
    • The Goal: A truly smart AI should be much more surprised by the reversed causal video than the reversed non-causal one.

The Big Reveal

The researchers tested 13 of the smartest video AI models available. Here is what they found:

  • The "Time" vs. "Logic" Gap: Many models were great at spotting that a video was backward (Level 1), but they failed to distinguish between a video where cause-and-effect was broken versus one where it wasn't (Level 2).
    • Analogy: It's like a student who knows the alphabet but can't read a sentence. They know the letters are out of order, but they don't understand the meaning.
  • Still Not Human: Even the best AI models were far behind human performance. Humans are naturally wired to understand cause and effect; the AIs are still just guessing based on patterns.
  • Bigger isn't Always Smarter (Yet): While bigger models and newer architectures generally did better, simply making the model bigger didn't automatically give it a "human-like" understanding of why things happen.

The Conclusion

YoCausal proves that just because an AI can generate a beautiful, realistic video, it doesn't mean it understands the world. It might be a master of "visual mimicry" but a novice at "logical reasoning."

The paper concludes that we are still far from building a true "World Model" (an AI that understands how the universe works). To get there, we need to teach these models not just to see the world, but to understand the rules that make the world tick.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →