← Latest papers
💬 NLP

Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision-Language Models

This paper introduces AoT-PsyPhyBENCH, a psychophysically validated benchmark demonstrating that current vision-language models struggle to determine the arrow of time in natural videos, revealing a fundamental gap in their temporal and causal reasoning capabilities compared to humans.

Original authors: Shiho Matta, Lis Kanashiro Pereira, Peitao Han, Fei Cheng, Shigeru Kitazawa

Published 2026-03-27
📖 4 min read☕ Coffee break read

Original authors: Shiho Matta, Lis Kanashiro Pereira, Peitao Han, Fei Cheng, Shigeru Kitazawa

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a magic video camera that can record the world. Now, imagine you play that video backwards.

If you see a shattered glass mug suddenly leap off the floor, pieces flying up to reassemble perfectly on the table, you know instantly: "That's backwards!" It defies the laws of physics. Gravity pulls things down, not up. Entropy (disorder) increases, it doesn't spontaneously decrease.

Humans have this superpower built into our brains. We don't need to think about it; we just know time flows one way.

This paper asks a very simple, yet terrifying question for Artificial Intelligence: Do modern AI video models have this same superpower?

The Experiment: The "Time-Travel" Test

The researchers created a benchmark called AoT-PsyPhyBENCH (Arrow of Time Psychophysics Bench). Think of it as a "driver's license test" for AI, but instead of driving a car, the AI has to drive a time machine.

They took hundreds of short, everyday videos (like a ball falling, an explosion, or someone building a sandcastle). They showed these clips to humans and to various AI models (like GPT-4, Gemini, and open-source models) and asked: "Is this playing forward or backward?"

The Results: The AI Got Lost in Time

The results were shocking.

  1. Humans are Time Masters: Humans got it right about 89% of the time. When a snowball fell up into the sky, we knew immediately it was reversed.
  2. AI is Time-Confused: The best AI models only got it right about 60% of the time. That's barely better than flipping a coin (50%).
    • The "Forward" Bias: Most AIs have a weird habit. If they are confused, they almost always guess "Forward." It's like a student taking a test who, when they don't know the answer, just circles "C" every time. They are so used to seeing videos play forward that they can't imagine the world running backward.
    • The "Thinking" Trap: The researchers tried to help the AI by asking it to "think step-by-step" (a technique called Chain-of-Thought). They hoped the AI would say, "Wait, the water is shrinking into a splash, that's impossible, so it must be backward."
      • What actually happened? The AI got worse. It would describe the scene perfectly ("I see a splash getting smaller") but then confidently conclude, "Therefore, this is playing forward." It was like a detective who found a clue but ignored it because they were too confident in their wrong theory.

Why Did the AI Fail?

The paper suggests the AI isn't "dumb" in the usual sense. It's great at recognizing objects. It knows what a "ball" is, what "snow" is, and what "falling" looks like.

However, the AI lacks Intuitive Physics.

  • Humans have a mental model of how the universe works (gravity, cause-and-effect). We know that if you drop a cup, it breaks. We know you can't un-break it.
  • AI only sees patterns. It has seen millions of videos of cups breaking. It knows "cup" + "falling" usually equals "broken." But it doesn't understand the rule that time only moves forward. It's like a parrot that can repeat the phrase "The cup fell" but doesn't understand why it fell or what would happen if you played the tape backward.

The Analogy: The Movie Buff vs. The Physicist

Imagine two people watching a movie played in reverse:

  • The AI (The Movie Buff): It sees a man walking backward. It thinks, "I've seen people walk backward in movies! This is normal!" It relies on what it has seen before.
  • The Human (The Physicist): It sees the man walking backward and thinks, "Wait, his feet are pushing off the ground in a way that defies gravity. The coffee cup is jumping from the floor to the table. This is impossible!" It relies on the laws of physics.

The AI is the Movie Buff. It has seen a lot of movies, but it hasn't learned the laws of the universe.

The Takeaway

This paper is a wake-up call. Even the most advanced AI models today are excellent at describing what they see, but they are terrible at understanding how the world works over time.

To build AI that truly understands the world, we can't just feed it more videos or ask it to "think harder." We need to teach it the fundamental rules of physics and causality—the invisible "arrow of time" that guides everything from falling apples to exploding fireworks. Until then, if you ask an AI to watch a video in reverse, it might just tell you it's a normal day.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →