MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
This paper introduces MultiRef-Compass, a unified benchmark comprising 350 curated samples and a four-dimensional evaluation protocol with 14 sub-metrics, designed to comprehensively assess the emerging capability of multi-reference-to-audio-video generation systems in preserving, binding, and composing multiple references into coherent synchronized content.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a director trying to film a movie scene. In the old days of AI video, you could only give the computer a simple script: "Make a video of a cat." The AI would guess what the cat looked like, where it was, and what sound it made. But now, creators want more control. They want to say, "Use this specific photo of my dog for the face, this other photo for his side profile, and this third photo for his favorite toy, all while he barks at a specific sound." This is the world of Multi-Reference-to-Audio-Video (MR2AV) generation. It's like asking an AI to be a master chef who must not only cook a meal but also use specific ingredients from three different jars, arrange them on the plate exactly as shown in a picture, and make sure the sizzling sound matches the cooking action perfectly. The problem is, we didn't have a good way to grade how well the AI was doing this complex dance. We had tests for simple recipes, but nothing for this multi-ingredient, multi-sensory challenge.
Enter MultiRef-Compass, a new "report card" designed by researchers to test exactly this skill. Think of it as a rigorous taste-test for AI video makers. The team built a special dataset of 350 carefully crafted scenarios—like a cooking show challenge where every contestant gets the exact same three photos of a person, a shoe, and a diner, plus a script saying, "The person tries on the shoe and asks if it fits." They then asked eight different AI models (including big names like Kling, Seedance, and Gemini) to generate the video.
The researchers didn't just look at whether the video looked "nice." They broke the grading down into four specific categories, like a detailed rubric:
- Basic Quality: Does the video look clear and sound good, or is it blurry and crackly?
- Reference Consistency: Did the AI actually use the photos you gave it? Did it keep the person's face the same, or did it accidentally swap them with a stranger? Did it glue the shoe onto the foot correctly, or did it just paste a sticker of a shoe on the screen?
- Audio-Visual Consistency: When the person in the video speaks, do their lips move? When they step on the floor, do you hear a footstep?
- Instruction Following: Did the AI do exactly what the script asked, or did it ignore the part about the shoe and just show the person walking?
To grade these, they used a mix of computer tools and a "super-smart AI judge" (a Large Multimodal Model) that acts like a film critic. They even added a special "re-judge" step to make sure the AI critic wasn't being too harsh or too easy, ensuring the scores were fair.
The results were a reality check. While some models are getting really good at making things look pretty, they are still struggling with the hard stuff. The paper suggests that current models often fail at binding—they might recognize the face in the photo but forget to attach it to the body, or they might mix up which person is supposed to be speaking. One model, Seedance 2.0, came out on top overall because it was the most balanced, doing well in all four categories. However, even the best models showed "substantial room for improvement." They found that when you ask an AI to juggle multiple references and complex instructions at once, it often drops the ball, creating videos where the characters look like cutouts, the voices don't match the speakers, or the story gets shuffled.
In short, MultiRef-Compass shows us that while AI video is getting impressive, it's not quite ready to be a reliable director for complex, multi-reference stories yet. It's a tool that needs more training to understand how to weave multiple clues together into a single, coherent, and natural-looking scene. The benchmark itself is now open for everyone to use, acting as a foundation to help future AI models learn how to be better at this specific, tricky kind of creative work.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.