What's the Catch? Evaluating Temporal Consistency in Vision-Language Models
The paper introduces TimeCatch, a benchmark revealing that while Vision-Language Models excel at detecting frame-level anomalies, they significantly struggle to identify temporal inconsistencies caused by frame swapping, indicating a fundamental gap in their ability to reason about temporal structure compared to humans.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the rapidly evolving world of artificial intelligence, a new generation of computer systems has emerged that can see and speak at the same time. These vision-language models are trained on vast amounts of images and text, allowing them to describe a photograph, answer questions about a scene, or identify objects with remarkable accuracy. As these systems grow more sophisticated, researchers are beginning to test them on moving images, hoping they can understand videos just as humans do. This is a crucial step for technology that interacts with the real world, such as self-driving cars that must predict the path of a pedestrian or robots that need to assemble parts in a specific order. However, a fundamental question remains: do these machines truly understand how time flows, or are they simply recognizing static pictures in a sequence?
To answer this, a team of researchers from the University of Copenhagen devised a simple but revealing test called TimeCatch. They wanted to see if these artificial intelligence models could spot when a video has been subtly broken. Imagine watching a short clip of a person pouring water into a glass. If the researchers swapped two consecutive frames so that the water appeared to jump backward or the glass filled before the cup was lifted, a human would instantly notice the error. The researchers created thousands of these "glitches" in both computer-generated animations and real-world footage, ranging from driving scenarios to competitive diving. They then asked the AI models to perform two tasks: first, to say whether a sequence contained a mistake, and second, to point out exactly where the mistake happened. They also included a control test where they replaced a single frame with static noise, a visual error that does not depend on the order of events, to see if the models could even see the frames at all.
The results revealed a striking gap between human perception and machine reasoning. When the models were shown a frame filled with static noise, they performed almost perfectly, identifying the corrupted image and locating it with near-ceiling accuracy. This proved that the models could see the individual pictures and were paying attention to the sequence. However, when the task shifted to finding the swapped frames that broke the flow of time, the models faltered. Their performance dropped to the level of random guessing. Even the most advanced models, which could correctly identify that a frame was missing or corrupted, failed to recognize that the order of events was impossible. In contrast, human participants in the study solved these temporal puzzles with high accuracy, easily spotting the illogical jumps in time.
The researchers explored whether this failure was due to the models being too small, the images being too similar, or the instructions being unclear. They tested models of different sizes, from smaller versions to massive ones with billions of parameters, and found that making the models larger did not solve the problem. They also checked if the models struggled because the swapped frames looked too much alike, but even when the visual differences were obvious, the models could not reason about the sequence. Changing the way the questions were asked or removing extra text descriptions did not help either. The study suggests that the issue is not a lack of visual power or processing capacity, but a fundamental inability to integrate information across time. These systems can describe a single moment in great detail, but they struggle to understand how one moment leads to the next.
This discovery highlights a significant limitation in current artificial intelligence. While these models are becoming incredibly good at describing what is visible in a single frame, they have not yet learned to reason about the continuity of events. For applications like autonomous driving or medical imaging, where understanding the progression of time is as important as recognizing objects, this gap is critical. The TimeCatch benchmark provides a clear, controlled way to measure this specific weakness, showing that despite their impressive abilities, today's vision-language models are still missing a key piece of the puzzle: the ability to truly understand the flow of time.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.