VGI-Bench: Probing Visual Intelligence in Video Generation Models
This paper introduces VGI-Bench, a comprehensive benchmark comprising 27 tasks and 810 instances designed to rigorously evaluate the visual reasoning capabilities of video generation models, revealing that even state-of-the-art systems like Seedance 2.0 achieve only 51.0% accuracy and struggle with reliable self-correction during the generation process.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a computer that doesn't just watch the world, but can imagine how it changes. For years, artificial intelligence has been trained to recognize static images: a cat sitting on a mat, a car parked on a street. More recently, these systems have learned to generate videos, creating moving pictures that look surprisingly real. Scientists now wonder if these video generators are doing something deeper than simply mimicking motion. They are asking if the machines are beginning to understand the rules of the physical world—gravity, cause and effect, and how objects interact—so that they can "think" through a sequence of events before showing the result. This question sits at the heart of a new effort to measure whether video-making AI can truly reason, or if it is merely guessing what a plausible ending looks like.
A team of researchers has built a new testing ground called VGI-Bench to answer this question. Instead of asking the AI to solve abstract puzzles drawn in simple lines, they fed it realistic photos of everyday scenes and asked it to generate videos that follow specific, logical rules. The tasks were designed to be challenging but possible, requiring the machine to simulate a process step-by-step. For example, an AI might be shown a photo of a toy car at the start of a maze and asked to generate a video of it driving to the exit without hitting the walls. Another task might show a flat, unfolded cardboard box and ask the AI to animate it folding itself into a 3D shape. The researchers did not just look at the final image; they scrutinized the entire video to see if the machine followed the rules at every moment, checking for things like objects passing through walls, cars teleporting, or items disappearing and reappearing.
When they tested the most advanced video generation models available today, the results were a mix of promise and limitation. The best-performing system, a model named Seedance 2.0, managed to solve only about half of the tasks correctly under these strict conditions. While the AI could often get the general idea right, it frequently failed to maintain a consistent story. In many videos, the physics broke down: a cup sitting on a tilted laptop would not roll off, or a car would drive straight through a solid wall. The machines also struggled with rules, sometimes skipping necessary steps or changing the identity of an object mid-way through the video. Even when the final result looked correct, the path to get there was often filled with impossible movements, suggesting the models are good at predicting a final state but poor at simulating the journey to get there.
The researchers dug deeper to understand why these failures happened. They found that the style of the input image mattered greatly; when the AI was given a realistic photo, it performed better than when given a simple line drawing, indicating that the models are trained on real-world images and struggle when the visual style changes. They also looked at how the AI "thinks" as it creates a video, step by step. They discovered that the models do not really correct their mistakes as they go. If the AI makes an error in the early stages of generating a video, it tends to stick with that error, refining the wrong idea rather than fixing it. The machine locks into a hypothesis early on and then polishes it, even if that hypothesis violates the laws of physics or the rules of the task.
This study suggests that while video generation models are becoming powerful tools for creating visual content, they have not yet developed the kind of reliable, step-by-step reasoning that humans use to navigate the world. They can simulate appearances and simple motions, but they lack a deep, consistent understanding of how things actually work over time. The new benchmark provides a clear way to measure this gap, showing that current systems are far from being true visual reasoners. By highlighting exactly where these models fail—whether in tracking objects, following rules, or correcting errors—the research offers a roadmap for building the next generation of artificial intelligence that can truly imagine the future.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.