← Latest papers
🤖 AI

RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement

This paper introduces RAVEN-Eval, a scalable, rubric-guided automated evaluation framework that leverages Large Multimodal Model (LMM) preference judgments to reliably distinguish fine-grained quality differences among state-of-the-art AI video generation models while minimizing human annotation costs.

Original authors: Ziheng Jia, Jiaying Qian, Zicheng Zhang, Xiaorong Zhu, Lancheng Gao, Xiongkuo Min

Published 2026-08-11
📖 4 min read☕ Coffee break read

Original authors: Ziheng Jia, Jiaying Qian, Zicheng Zhang, Xiaorong Zhu, Lancheng Gao, Xiongkuo Min

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are the director of a massive, high-tech movie studio where the actors are actually super-smart computers. These computers, known as AI video generators, can now create entire scenes from scratch just by reading a sentence you type. But here's the catch: they are getting so good that the difference between a "pretty good" video and a "mind-blowing" one is as subtle as the difference between two shades of blue. In the past, you could just ask a human to watch and say, "That one looks better." But now, with thousands of videos being made every day, hiring enough humans to watch them all would cost a fortune and take forever. Plus, even humans get tired and might miss tiny details. This is the problem scientists are facing: how do you grade these super-advanced computer artists when they are all performing at a near-perfect level, and you can't afford to have a human watch every single clip?

This is where the paper "RAVEN-Eval" comes in. It's like a new, ultra-strict referee system designed to judge these AI movies without needing a human to sit through every single one. The researchers realized that instead of asking a computer, "Rate this video from 1 to 10," which is often confusing and inconsistent, it's better to ask, "Which of these two videos is better?" This is called a "pairwise comparison," and it's much easier for both humans and computers to decide. But to make sure the computer referee doesn't just pick the flashiest video, the team gave it a special rulebook, or "rubric," for every single task. It's like giving a judge a checklist that says, "If the prompt asked for a red ball, the ball must be red," or "If the prompt asked for a hydraulic press crushing a duck, the duck must squish, not explode." By using these specific rules, the system can spot the tiny, fine-grained differences that separate the best AI models from the rest.

The team behind this study, led by researchers from Shanghai Jiao Tong University and the Shanghai Artificial Intelligence Laboratory, built a massive testing ground called the RAVEN-Eval Benchmark. They didn't just throw random prompts at the AI; they carefully curated 250 different tasks, ranging from simple text-to-video requests to complex image-to-video challenges where the AI has to fill in the missing middle of a scene. They generated over 4,500 videos to test 20 of the world's top AI video models. Instead of letting the AI judge itself blindly, they used a "judge" system powered by Large Multimodal Models (LMMs)—which are basically AI brains that can see and understand images and text. These AI judges were given the specific rulebooks for each task and asked to compare videos side-by-side.

The results were fascinating. The researchers found that when they gave the AI judges these detailed, task-specific rulebooks, the system became incredibly good at telling the difference between top-tier models. In fact, their new system agreed with human experts much more closely than previous methods that just asked for a general score or used fixed checklists. They even created two leaderboards: one to rank the 20 AI video models (showing who is currently the best artist) and another to rank the 13 different AI judges themselves (showing which AI brain is the most accurate referee). One of the coolest tricks they invented is called "anchor-based insertion." Imagine you have a leaderboard of the top 20 runners, and a new runner shows up. Instead of making the new runner race against all 20 people (which would take forever), you just have them race against a few carefully chosen "anchor" runners from the top, middle, and bottom of the pack. This lets you figure out where the new runner fits in the ranking with much less effort.

Ultimately, the paper suggests that this new, rule-guided way of using AI to judge AI is a scalable and reliable path forward. It proves that we don't need to rely solely on expensive human annotators to keep up with the rapid pace of AI video generation. By using smart, rule-based comparisons, we can build a trustworthy system that evolves as fast as the models it is trying to measure, ensuring that as these digital artists get better, our ability to appreciate and rank their work gets better, too.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →