AgenticVBench: Can AI Agents Complete Real-World Post-Production Tasks?
The paper introduces AgenticVBench, a benchmark of 100 real-world video post-production tasks constructed with industry expert input, which reveals that current frontier multimodal AI agents achieve significantly lower success rates than human experts and are highly sensitive to the choice of execution harness.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to be a movie editor. You don't just want it to answer questions about a movie; you want it to actually make one. It needs to cut scenes, fix bad audio, rearrange shuffled clips, and turn a three-hour documentary into a snappy 60-second social media video.
This paper, AgenticVBench, is like a final exam for AI robots trying to learn this job. The creators built a test based on real-world work done by 20 human movie experts. Here is the breakdown of what they did and what they found, using simple analogies.
1. The Test: "The Movie Editor's Gauntlet"
The researchers created 100 specific tasks divided into four categories, like four different stations in a movie editing workshop:
- Assembly (The Puzzle): Imagine you have a storyboard (a comic book sketch of a movie) and a pile of video clips. Your job is to pick the exact right clip for every drawing in the sketch. The AI has to understand camera angles and shot sizes, not just "does this look like a dog?"
- Repair (The Doctor): You are given a video with a hidden "sickness" (like a weird echo, a color glitch, or a scene played in the wrong order). The AI must find the sickness, fix it, and write a report explaining what it did.
- Sequencing (The Time Traveler): You are given a pile of video clips that have been shuffled out of order, along with a short summary of the story. The AI must put the clips back in the correct chronological order so the story makes sense.
- Repurpose (The Shrink): You have a long movie (maybe 3 hours) and a request to make a short trailer or a recap. The AI must watch the whole thing, decide what to keep, what to cut, and assemble a new, short video that fits specific rules (like "must be 60 seconds" or "must be funny").
2. The Players: AI vs. Humans
The researchers tested 7 of the smartest AI "brains" (Vision-Language Models) available today. They ran these brains through two types of "bodies" (harnesses):
- Native Bodies: The official tools the AI companies built for their own models (like a custom-built car).
- Open-Source Bodies: Tools built by the community that anyone can use (like a generic car chassis).
They also had 20 human experts do the same tasks to set a "gold standard" for how well a human should do.
3. The Results: The AI is Still a Rookie
The results were a bit of a reality check for the AI world:
- The Score: The best AI team managed to get about 31% of the tasks right.
- The Gap: Human experts scored much higher (around 70–95% depending on the task). The AI is currently about 43 to 65 percentage points behind a human.
- The "Body" Matters: It wasn't just about how smart the AI brain was. The "body" (the software tools wrapping the AI) mattered just as much. Sometimes, the same AI brain got a 38% score with one toolset and only an 18% score with another. It's like having a Formula 1 driver in a minivan; the driver is great, but the car limits them.
4. Where the AI Stumbled
The paper looked closely at why the AI failed, finding four main "bad habits":
- Getting Lost in the Library (Long-Context Loss): For the "Repurpose" task (making a short video from a long one), the AI would get so busy trying to read the entire transcript of the movie that it ran out of time and never actually made the video. It was like a student reading the whole encyclopedia instead of writing the essay.
- Bad Timing (Temporal Reasoning): For the "Repair" task, the AI often knew what was wrong but couldn't figure out when it happened. It might fix the right sound effect but apply it to the wrong scene, shifting the timing by 15 to 100 seconds.
- Wrong Tools (Modality Misalignment): The AI would try to fix a visual problem with audio tools, or vice versa.
- Hallucinations: Sometimes the AI would invent a scene that didn't exist or claim a defect was fixed when it wasn't.
5. The Big Takeaway
The paper concludes that while AI is getting better at understanding images and text, it is still very far from being able to plan and execute complex, multi-step creative jobs like movie editing.
It's not just about making the AI "smarter." The paper argues we also need to build better "toolboxes" (harnesses) that help the AI manage its time, check its own work, and understand how to use video editing software properly. Until we fix both the brain and the toolbox, AI movie editors will remain assistants that need heavy human supervision, not the directors themselves.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.