← Latest papers
💻 computer science

VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation

This paper introduces VA-Judger, a chain-of-thought reward model trained on a new large-scale human preference dataset (VAPref-10K) and a novel benchmark (VA-Judger-Bench) to overcome the limitations of existing separate quality metrics, thereby providing more coherent and human-aligned rewards that significantly improve joint video-audio generation through reinforcement learning.

Original authors: Yinming Huang, Shuyuan Tu, Xi Yan, Zihan Yang, Jianhua Han, Xu Hang, Yu-Gang Jiang, Zuxuan Wu

Published 2026-08-20
📖 4 min read☕ Coffee break read

Original authors: Yinming Huang, Shuyuan Tu, Xi Yan, Zihan Yang, Jianhua Han, Xu Hang, Yu-Gang Jiang, Zuxuan Wu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where a computer can not only paint a moving picture but also compose the exact sounds that belong to it: the rustle of leaves, the crackle of a fire, or the specific timbre of a voice speaking in a room. This is the frontier of joint video-audio generation, a field where artificial intelligence learns to create visual scenes and synchronized soundscapes simultaneously. For years, researchers have struggled to teach these systems to produce content that feels truly real to human viewers. The problem is not just making a clear image or a crisp sound; it is ensuring that the two elements work together as a single, coherent experience that matches the story being told. When a cow in a video attempts to play a guitar, the sound must be a clumsy, mooing strum, not a perfect musical chord. If the computer gets the visual details right but the sound wrong, or if the audio and video are slightly out of step, the illusion breaks. Until now, the tools used to judge these creations have been like a panel of inspectors checking the paint, the engine, and the wheels of a car separately, missing the fact that the car itself might not drive at all.

A team of researchers has introduced a new approach to solve this disconnect. They built a system called VA-Judger, designed to act as a critical eye and ear for artificial intelligence. Instead of relying on separate, rigid rules to measure video quality, audio quality, and timing, this new system learns to evaluate the entire package the way a human does. The researchers started by creating a massive library of comparisons, gathering thousands of pairs of video-audio clips generated by different computer models. For each pair, human annotators watched and listened, deciding which one felt better and explaining exactly why. They noted whether one clip followed the written story more closely, if the lips moved in time with the words, or if the background noise felt natural. This collection of human choices became the training ground for VA-Judger.

The training process was carefully structured to teach the system how to think. First, the model learned on easy examples where the difference between a good clip and a bad one was obvious, much like a student learning to distinguish between a clear photo and a blurry one. Once it mastered the format of giving a structured critique, the researchers presented it with much harder cases where two clips were nearly identical in quality. Here, the model had to learn to spot subtle flaws, such as a slight delay in a sound effect or a gesture that didn't quite match the emotion of the voice. To ensure the model was learning from human truth rather than just mimicking patterns, the researchers used a filtering step where human experts verified the model's choices on these difficult pairs, keeping only the reasoning that aligned with human judgment. Finally, the system was refined through a process of trial and error, where it received specific feedback on each part of its reasoning, encouraging it to develop a deeper understanding of why one video-audio pair succeeded while another failed.

The results of this work show that the new system aligns far more closely with human preference than previous methods. When tested against a wide range of existing tools that measure video and audio separately, VA-Judger proved much better at predicting which clip a person would actually enjoy. In one set of tests involving hundreds of comparisons, the new system agreed with human choices significantly more often than the old metrics, which often got confused by clips that looked good but sounded wrong, or vice versa. More importantly, when the researchers used VA-Judger to help train a video-generation model, the output improved dramatically. The new model produced clips that were not only sharper and clearer but also felt more complete and faithful to the original story. In head-to-head comparisons, human viewers chose the clips generated with this new guidance more than six times as often as those from the untrained base model, and more than twice as often as those from a model trained with older methods.

This work demonstrates that for artificial intelligence to create truly convincing video and audio, it cannot be judged by a checklist of isolated features. The quality of a generated scene depends on the seamless integration of sight, sound, and story. By teaching the computer to evaluate these elements together, using the nuanced feedback of human observers, the researchers have provided a path toward generating content that feels less like a simulation and more like a genuine moment captured in time. The findings suggest that the future of creative AI lies not in better individual sensors, but in systems that understand the whole picture.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →