VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality
This paper introduces VGA-BenchV2, an expanded human-aligned benchmark and multi-model framework that significantly scales up annotations and evaluation dimensions to jointly assess video generation quality and aesthetics, while also establishing an evaluation-to-optimization pipeline to fine-tune generators for improved aesthetic alignment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the last few years, computers have learned to create moving pictures from simple written descriptions. This technology, known as text-to-video generation, has moved from experimental sketches to tools capable of producing coherent, stable, and visually striking scenes. Artists, filmmakers, and advertisers are beginning to use these systems to imagine new worlds, but a fundamental question remains: how do we know if the result is actually good? For a long time, scientists measured these videos by checking technical details, such as whether the frames flowed smoothly or if the computer followed the instructions literally. However, a video can be technically perfect and still feel flat, lifeless, or artistically dull. The missing piece has been a way to measure the human experience of watching the video—its beauty, its mood, and its emotional resonance.
A team of researchers from Ant Group, the Beijing Film Academy, and the Beijing Institute for General Artificial Intelligence has tackled this challenge with a new framework called VGA-BenchV2. Their work moves beyond simple technical checks to create a system that understands and evaluates video aesthetics the way a human does. They built a massive library of over 60,000 videos generated by 12 different computer models, using 1,016 carefully written prompts to test specific visual qualities. To teach their computers how to judge these videos, the team enlisted human experts to provide over 36,000 new ratings and labels. This human feedback was used to train a hybrid team of artificial intelligence evaluators: one specialized in giving a continuous score for beauty, and two others that act like visual critics, identifying specific artistic tags and checking if the video makes sense. The result is a system that not only ranks how good a video looks but also uses that judgment to help the computer models learn and improve themselves.
The core of this new framework is a shift from passive measurement to active guidance. In the past, benchmarks were like report cards that simply told a student they passed or failed. VGA-BenchV2 acts more like a tutor that explains exactly what needs improvement and then helps the student practice. The researchers organized their evaluation into two main categories: aesthetics and generation quality. The aesthetic side looks at the artistic elements, such as lighting, color harmony, composition, and the overall mood of the scene. The generation side checks if the video is physically plausible, if the characters move naturally, and if the story matches the written prompt. By breaking these broad concepts down into 52 specific sub-categories, the team ensured that the evaluation was detailed and precise, covering everything from the type of camera shot to the realism of fluid motion.
To build this system, the team started with a large set of prompts, which are the written instructions given to the video generators. These 1,016 prompts were designed to trigger specific visual effects, ensuring that the resulting videos could be judged fairly against clear criteria. They then collected the output from 12 leading video generation models, creating a pool of more than 60,000 videos. The crucial step was the human annotation phase. The researchers gathered a massive amount of human feedback, adding 36,000 new task-level annotations to their existing data. This included over 16,000 ratings for aesthetic quality, where humans scored videos on a scale of zero to ten, and over 13,000 labels for aesthetic tagging, where humans identified specific visual attributes like the color of the light or the type of shot. They also added thousands of notes on generation quality to ensure the videos were realistic and consistent.
With this vast amount of human-labeled data, the team trained three specialized artificial intelligence evaluators. The first, called VAQA-Net, learned to predict a continuous score for how beautiful a video is, effectively mimicking the human eye for composition and tone. The other two evaluators, VTag-Net and VGQA-Net, are based on large vision-language models, which are advanced systems capable of understanding both images and text. These models were trained to act as critics: one identifies specific artistic tags, such as whether a scene uses soft lighting or high contrast, while the other answers detailed questions about the video's realism and consistency. This combination allows the system to provide both a numerical score and a detailed, interpretable explanation of why a video is good or bad.
The true power of VGA-BenchV2 lies in its ability to close the loop between evaluation and improvement. Once the evaluators were trained to match human preferences, the researchers used them as a reward system for reinforcement learning. In this process, the computer models that generate the videos are shown their own creations and given a score based on the aesthetic quality predicted by the evaluator. If the score is high, the model is encouraged to make similar videos; if the score is low, it is nudged to change its approach. In a specific test, the researchers applied this method to fine-tune a video generator called Wan2.1. The results showed that the model's average aesthetic score increased, and the resulting videos displayed more appealing visual compositions and higher overall quality. This demonstrates that the framework does not just measure performance; it actively drives the technology toward better artistic outcomes.
The study also provided a comprehensive ranking of the current state of the art in video generation. By applying their new evaluators to the 12 different models, the team was able to compare them across all 52 dimensions. The results revealed significant differences in how well each model handled artistic control versus technical consistency. Some models excelled at following complex prompts, while others produced more visually pleasing results. The data showed that while many models have improved in technical fidelity, there is still a wide gap in their ability to consistently produce videos that align with human aesthetic preferences. This highlights the importance of the new framework, which provides a standardized way to measure these subtle but critical qualities.
Ultimately, VGA-BenchV2 represents a significant step forward in making artificial intelligence more attuned to human sensibilities. By combining a massive dataset of human judgments with a sophisticated hybrid evaluation system, the researchers have created a tool that can assess video generation with a depth and nuance that was previously impossible. The framework bridges the gap between technical capability and artistic expression, offering a path for future models to learn not just how to generate pixels, but how to create art that resonates with people. As video generation continues to evolve, tools like this will be essential for ensuring that the machines we build can create content that is not only real but also beautiful.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.