← Latest papers
💻 computer science

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward

This paper introduces AVCap, a comprehensive framework comprising the high-quality AVCap-100K dataset, a Detail-Aware GRPO-optimized model, and a specialized benchmark with atomic-level metrics, to overcome existing limitations in detailed audio-video joint captioning.

Original authors: Mingyang Wu, Kaituo Feng, Bohao Li, Kaixiong Gong, Zihao Yin, Xiangyu Yue

Published 2026-08-10
📖 4 min read☕ Coffee break read

Original authors: Mingyang Wu, Kaituo Feng, Bohao Li, Kaixiong Gong, Zihao Yin, Xiangyu Yue

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to watch a movie and tell you exactly what happens. In the world of artificial intelligence, this is called "video captioning." For a long time, these robots were like people who only had their eyes open; they could describe the colors and movements on the screen but were completely deaf to the soundtrack. They missed the scary creak of a floorboard, the specific accent of a character, or the way the music swelled to match a dramatic moment. Recently, scientists have built "multimodal" models that can both see and hear, but teaching them to describe a video with perfect detail is like trying to write a recipe for a complex dish while blindfolded and wearing earplugs. The robots often guess wrong, mixing up what they see with what they think they hear, or they skip the tiny, important details that make a story real. This paper tackles that problem by asking: How do we teach a robot to be a super-observant, detail-obsessed narrator who never misses a beat, a sound, or a glance?

The researchers behind this study, AVCap, decided that the old way of teaching these robots wasn't working because the "homework" they were given was too vague and the "grading" was too easy. To fix this, they built three new tools: a massive, ultra-detailed library of video descriptions, a new way to grade the robots' work that punishes tiny mistakes, and a special test to see if the robots are actually paying attention to the fine print.

First, they realized that existing video libraries were like a collection of blurry snapshots; they had the main events but lacked the texture. So, the team created AVCap-100K, a dataset containing 100,000 video clips paired with incredibly rich descriptions. Imagine taking a 60-second video and breaking it down not just into "a dog runs," but into "a golden retriever with a muddy paw pads sprints across the grass, its collar jingling, while a distant lawnmower hums." They achieved this by using a clever pipeline that first listens to the audio and watches the video separately to get the facts straight, and then combines them into one perfect story. This ensures the robot doesn't hallucinate (make things up) by guessing that a sound belongs to a visual event that isn't actually there.

Next, they tackled the problem of how to teach the robot to be this detailed. Previous methods were like a teacher who only gave a grade based on whether the story "sounded good" overall. If the robot missed a specific sound effect or got a color wrong, the teacher might not notice. The authors introduced a new training method called Detail-Aware GRPO (Da-GRPO). Think of this as a strict, hyper-focused editor. Instead of just giving a score, this editor acts like a detective. It takes the robot's description and asks a series of specific questions based on the real video, such as "What color was the shirt?" or "What sound did the door make?" If the robot's answer doesn't match the facts, it gets a penalty. This forces the robot to learn that getting the tiny, atomic details right is just as important as telling the main story.

Finally, to prove their new robot was actually better, they built AVCap-Bench and a new scoring system called AVCap-Score. Instead of just checking if the robot used the right words, this system checks if the robot actually knows the facts. It's like a pop quiz where the robot has to answer 20 specific questions about the video it just described. If it can't answer them correctly, it loses points, no matter how well-written the story was.

The results were impressive. Their new model, AVCap, trained with this strict "detective" method, became a champion at describing videos. On standard tests, it outperformed many other open-source models and even matched or beat some of the most expensive, commercial "black box" models used by big tech companies. For instance, on a specific test called AVCap-Bench, their 30-billion-parameter model scored 56.94, while the best commercial model they compared it to, Gemini-2.5-Pro, scored lower. They also showed that their model could describe videos with fewer "hallucinations" (mistakes where it invents details) and fewer missed events.

In short, this paper suggests that if you want a robot to truly understand a video, you can't just let it guess; you have to give it a massive library of detailed examples and then grade it like a strict teacher who checks every single fact. By doing this, they created a system that doesn't just summarize a video but captures the full, rich experience of seeing and hearing it, paving the way for robots that can tell stories with the precision of a human observer.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →