← Latest papers
🤖 AI

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

This paper introduces Sci-VBench, a comprehensive benchmark comprising 1,253 expert-annotated examples across four scientific disciplines to evaluate the knowledge- and reasoning-intensive video generation capabilities of 16 frontier models, revealing that while visual realism has improved, significant gaps remain in scientific and causal correctness, particularly between proprietary and open-source systems.

Original authors: Diandian Zhang, Tingyu Song, Lin Fu, Zheyuan Yang, Yilun Zhao

Published 2026-08-11
📖 6 min read🧠 Deep dive

Original authors: Diandian Zhang, Tingyu Song, Lin Fu, Zheyuan Yang, Yilun Zhao

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you're watching a movie where a wizard casts a spell, a rocket launches, or a heart beats. For years, the biggest question for computer scientists was simple: "Does this look real?" If the pixels were sharp, the colors bright, and the motion smooth, we cheered. But there's a new, much trickier question now: "Does this actually make sense?"

This is the world of video generation, where artificial intelligence (AI) tries to create moving pictures from just a sentence of text. Think of it like a super-creative robot that reads a recipe and tries to cook the dish. Until recently, we were mostly checking if the food looked tasty. But what if the robot accidentally put salt in the cake instead of sugar? It might look like a perfect cake, but it would taste terrible and fail the test. In science, this is dangerous. If an AI simulates a chemical reaction or a medical procedure, it can't just look good; it has to follow the strict, invisible laws of physics and biology. If the AI breaks those laws, the video is a lie, no matter how pretty it is.

This is exactly the problem a new paper tackles. The researchers built a massive "exam" called Sci-VBench to test if AI video makers can actually understand science, not just mimic it. They didn't just ask, "Is the video pretty?" They asked, "Did the robot actually understand how the world works?"

The Great Science Video Exam

The team created a giant test bank with 1,253 different challenges, covering 60 different subjects like physics, medicine, engineering, and even history. Imagine a test where you have to generate a video of a specific experiment, like hitting a knee with a hammer to see the leg kick, or watching two balls roll down different tracks to see which wins. The catch? The AI has to figure out the hidden rules of how those things move on its own. It can't just guess; it has to get the science right.

To grade these videos, the researchers didn't just rely on a computer's opinion. They brought in 61 human experts—real scientists and students from universities—to act as the judges. These experts wrote down exactly what a "perfect" video should look like, step-by-step, creating a strict scoring guide (a rubric) for every single test. They checked four things:

  1. Does it look good? (Is the video clear and smooth?)
  2. Did it listen? (Did it include all the specific things the prompt asked for?)
  3. Is the science right? (Did the objects move according to real laws of nature?)
  4. Is the story consistent? (Did the video make sense from start to finish without glitching?)

The Shocking Results: Pretty but Clueless

When they ran 16 of the world's most advanced AI video models through this exam, the results were a bit of a wake-up call.

First, the good news: The AI models are getting incredibly good at making things look real. When the judges checked for basic visual quality (like smoothness and color), almost all the models scored very high, clustering tightly together. They are all masters of the "painting" part of the job.

But then came the bad news. When the judges checked if the AI actually understood the science, the scores dropped dramatically and varied wildly.

  • The proprietary models (the big, expensive ones from companies like Google, OpenAI, and Alibaba) generally did better. The top performer, Gemini-Omni-Flash, managed to get the science right most of the time, but even it wasn't perfect.
  • The open-source models (the free ones anyone can download) struggled significantly more. While they looked just as pretty as the expensive ones, they often failed the science tests. For example, in a test about a knee-jerk reflex, some models made the wrong leg kick, or made the leg move in a way that defies human anatomy.

The paper found a huge gap between "looking real" and "being real." The AI can create a video of a ball rolling down a hill that looks stunning, but if the ball suddenly floats up or changes color in the middle of the roll, the AI has failed the science test. The researchers noted that this gap is especially wide in Scientific and Causal Correctness—the ability to understand cause and effect.

Why the AI Gets It Wrong

The researchers dug deeper to see why the AI fails. They found three main types of mistakes:

  1. Ignoring Instructions: The AI might miss tiny details, like forgetting to show a specific tool mentioned in the prompt.
  2. Fake Science: This is the big one. The AI prioritizes making the video look "cool" or "smooth" over making it physically accurate. It might make a chemical reaction happen instantly because it looks dramatic, even though real chemistry takes time.
  3. Time Travel Glitches: The video might start with two balls, but halfway through, they merge into one, or the lighting changes randomly. The AI loses track of the story as time passes.

Interestingly, the researchers tried a simple fix: they rewrote the prompts to be more explicit, telling the AI exactly what to do. This helped a little bit, improving the scores for some models, but it didn't fix the core problem. The AI still couldn't generate the deep, mechanical understanding of how the world works. It's like giving a robot a more detailed recipe; it might follow the steps better, but if it doesn't understand why you mix the ingredients, it still won't bake a good cake.

The Bottom Line

The paper concludes that while AI video generators have become amazing artists, they are still clumsy scientists. They can paint a beautiful picture of a rocket launch, but they often don't understand the physics of the launch itself. The gap between "looking real" and "being real" is still wide, and until AI can truly grasp the rules of the universe, its scientific videos will remain just that—illusions. The researchers suggest that to move forward, we need to stop just asking "Is it pretty?" and start demanding, "Does it make sense?"

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →