LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
This paper introduces LongAV-Compass, a systematic benchmark designed to unify the evaluation of minute-scale audio-visual generation across text, image, and video conditioning modalities by employing a comprehensive framework that assesses over 20 fine-grained dimensions of quality, consistency, and coherence.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a film director. For a long time, the movie industry only asked you to make 5-second trailers. You could make those look amazing, and there were plenty of judges to tell you if the lighting was good or if the actor smiled at the right time.
But now, the industry is asking for full 60-minute movies. Suddenly, making a single good scene isn't enough. You have to make sure the main character looks the same in the final scene as they did in the first, the story makes sense from start to finish, and the sound effects match the action perfectly.
The problem? The judges (evaluation tools) are still only trained to grade 5-second trailers. They don't know how to spot a plot hole that happens 40 minutes in, or when a character's face starts to melt because the movie got too long.
Enter "LongAV-Compass."
This paper introduces a new, specialized "Judge" designed specifically to grade these new, long-minute videos. Here is how it works, broken down simply:
1. The Three Types of Directors (The Tasks)
The new judge tests three different ways of making a movie:
- Text-to-Movie (T2AV): You give the judge a script (text), and it has to make the whole movie.
- Photo-to-Movie (I2AV): You give the judge a single photo of a character, and it has to make them move and act out a story while keeping their face looking exactly like the photo.
- Video-to-Movie (V2AV): You give the judge the first 15 seconds of a movie, and it has to finish the rest of the story without changing the style or the characters.
2. The "Compass" Map (The Benchmark)
The researchers built a library of 284 specific test cases. Think of these as 284 different "scripts" or "challenges" ranging from simple vlogs to complex commercial ads.
- They organized these by difficulty: Some are easy (one person talking), while others are hard (a car chase with multiple people and specific physics).
- They cover different genres: Like "Brand Ads" (selling a product) or "Content Creator" (telling a funny story).
3. How the Judge Grades (The Metrics)
Instead of just giving a single score like "8 out of 10," the LongAV-Compass acts like a diagnostic doctor. It checks the video on over 20 different "vital signs":
- Did the story happen? (Event Fulfillment): Did the video actually do what the script asked?
- Did the actor stay the same? (Identity Consistency): Did the main character's face or clothes change weirdly halfway through?
- Is the story smooth? (Continuity): Did the video jump around or freeze between scenes?
- Do the sounds match? (Synchronization): When a door slams in the video, does the slam sound happen at the exact same time?
- Is the whole movie watchable? (Holistic Presentation): If you watched the whole thing, would it feel like a complete, polished product?
4. The Results (Who Passed the Test?)
The researchers tested 11 different AI video generators (a mix of big commercial companies and open-source projects) using this new Compass.
- The Big Winners: The top commercial models (like "Seedance" and "Kling") generally did the best. They could keep characters consistent and tell a coherent story for a full minute.
- The Struggles: Many open-source models could make a pretty picture for a few seconds, but as the video got longer, the stories fell apart, characters changed faces, or the audio got weird.
- The "Product" Problem: The hardest test was "Brand Ads" (selling a product). Most AIs struggled here. They couldn't reliably show a product being used step-by-step without messing up the logic or the visuals.
5. Why This Matters
The paper argues that we can't just look at a "Leaderboard" with one number anymore. An AI might be great at making a 5-second clip but terrible at a 60-minute one.
LongAV-Compass is a tool that helps us see exactly where these AI systems are failing. It's like a mechanic's diagnostic computer that tells you, "The engine runs fine for 10 minutes, but the brakes fail after 20," rather than just saying, "This car is broken."
In short: The world of AI video is growing up from making short clips to making full movies. This paper built the first ruler that can actually measure if those movies are any good.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.