← Latest papers
💻 computer science

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

This paper introduces AVE-Compass, a comprehensive benchmark designed to holistically evaluate audio-video editing capabilities through fine-grained checklist-based metrics, and proposes AVE-Agent, a modular framework that leverages self-reflection and iterative feedback to significantly improve cross-modal instruction following and fidelity in complex editing tasks.

Original authors: Yuqing Wen, Yukai Huang, Qianqian Xie, Jiangtao Wu, Yibin Lin, Yikai Gu, Jialu Chen, Yuanxing Zhang, Jiaheng Liu

Published 2026-07-29
📖 3 min read☕ Coffee break read

Original authors: Yuqing Wen, Yukai Huang, Qianqian Xie, Jiangtao Wu, Yibin Lin, Yikai Gu, Jialu Chen, Yuanxing Zhang, Jiaheng Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a director trying to edit a movie. In the old days, computers were like clumsy assistants who could only touch the picture. If you asked them to change a sunny sky to a storm, they would darken the clouds but leave the sound of birds chirping happily in the background. It looked weird and felt wrong because, in the real world, sight and sound are best friends; they hold hands and move together. This field of science is called multimodal editing, where we try to teach computers to understand that changing a visual action often requires changing the sound, and vice versa. The big question researchers are asking is: "Can we build a computer editor that doesn't just swap pixels or noise, but actually understands the whole scene so it doesn't break the magic?"

Enter AVE-Compass, a new "report card" created by researchers to grade how well these computer editors are doing. Think of it like a super-strict film critic who doesn't just watch the final movie; they have a checklist of 2,688 tiny details to inspect. They check if the editor followed your instructions, if they accidentally deleted the background music you wanted to keep, and if the new rain sounds actually match the new stormy sky. The researchers found that even the smartest current computers are still struggling. They often get the visual part right but mess up the audio, or they change the sound so much that it no longer fits the picture. It's like a musician who plays the right notes but in the wrong key, making the whole song feel off.

To fix this, the team built a new helper robot called AVE-Agent. Instead of just blindly trying to edit the video, this agent acts like a careful project manager. It breaks your big, messy request (like "make it look like a heavy rainstorm") into a list of small, logical steps. It plans ahead, checks its own work after every step, and if it makes a mistake, it says, "Oops, let me try that again," before moving on. When they tested this new agent, it did a much better job than the others. It followed instructions more closely, kept the original sounds and sights that shouldn't change, and made the final result look and sound more natural. While it's not perfect yet, this new approach suggests that if we teach computers to plan and check their work step-by-step, they might finally become the reliable co-directors we've been waiting for.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →