← Latest papers
💻 computer science

DirectorBench: Diagnosing Long-Form Video Generation with Personalized Multi-Agent Evaluation

DirectorBench introduces a personalized multi-agent diagnostic benchmark that evaluates long-form video generation across five dimensions and multiple user profiles to localize specific workflow failures and align more closely with human judgment than traditional aggregate scoring methods.

Original authors: Jiamin Chen, Qianben Chen, Jiawen Zhang, Yidi Wu, Yuchen Li, Xiaokun Zhang, Wangchunshu Zhou, Chen Ma

Published 2026-05-29
📖 4 min read☕ Coffee break read

Original authors: Jiamin Chen, Qianben Chen, Jiawen Zhang, Yidi Wu, Yuchen Li, Xiaokun Zhang, Wangchunshu Zhou, Chen Ma

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a movie director. In the past, AI video generators were like talented but short-attention-span actors who could only perform a single, perfect 5-second scene. If you asked for a whole movie, the AI would just stitch those scenes together clumsily, resulting in a jumbled mess where the characters changed clothes, the story made no sense, and the audio was out of sync.

DirectorBench is a new "quality control" system designed specifically to grade these long, complex AI movies. Instead of just giving a single grade (like "B-"), it acts like a specialized film critic panel that breaks the movie down into tiny pieces to find exactly where it went wrong.

Here is how it works, using simple analogies:

1. The Problem: The "One-Size-Fits-All" Grade is Broken

Old ways of testing AI videos were like grading a student's essay based only on their handwriting. They looked at if the pictures were clear or if the video didn't freeze. But for a long movie, you also need to check: Did the story make sense? Did the character's voice match their lips? Did the music fit the mood?

Furthermore, everyone has different tastes. A horror fan wants scary sounds; a romance fan wants emotional music. Old tests gave one single score to everyone, ignoring that a video might be perfect for one person but terrible for another.

2. The Solution: The "Personalized Critic Panel"

DirectorBench solves this by using a team of AI agents (specialized robots) that act as different types of critics.

  • The Script Agent: Reads the story to see if the plot makes sense.
  • The Visual Agent: Checks if the lighting, camera angles, and character faces stay consistent.
  • The Audio Agent: Listens to the music and voiceovers.
  • The Sync Agent: Checks if the lips move with the words and if the sound matches the action.

Instead of just saying "This movie is 7/10," this panel gives a diagnostic report. It might say: "The story is great, the visuals are pretty, but the transition between Scene 2 and Scene 3 is a disaster." This helps the creators know exactly what to fix.

3. The "User Profile" Twist

Imagine you are ordering a custom cake.

  • Profile A (The Storyteller): Cares mostly about the plot. They don't mind if the frosting is a little messy, as long as the story is good.
  • Profile B (The Visual Artist): Cares mostly about how the cake looks. They might forgive a boring story if the design is stunning.

DirectorBench lets you plug in these different "User Profiles." It re-runs the evaluation for each profile. A video might get a low score for the Storyteller but a high score for the Visual Artist. This proves that quality depends on who is watching.

4. What They Discovered (The "Bottlenecks")

The researchers tested several different AI movie-making systems using DirectorBench and found some surprising truths:

  • The "Seam" Problem: The biggest failure point isn't inside a single scene; it's between scenes. The AI is great at making one shot, but terrible at connecting two shots smoothly. It's like a musician who can play a perfect note but can't play a smooth melody. The "transition quality" was the lowest score across all systems.
  • The "Brain" vs. The "Hands": They tested different "brains" (the AI models that plan the movie) while keeping the "hands" (the tools that actually draw the video) the same. They found that changing the brain mostly changed the story and the coordination of the movie, but didn't fix the basic visual glitches. The "hands" (the video generation tools) were the limiting factor for visual quality, regardless of how smart the planner was.
  • The "Average" Lie: When they used a generic, "average" score, they missed huge differences. A video that looked "okay" to a neutral observer might be a total failure for a specific user type.

Summary

DirectorBench is a new tool that stops treating AI video generation like a simple math problem with one right answer. Instead, it treats it like filmmaking: it checks the story, the visuals, the sound, and the transitions separately, and it understands that a "good" movie depends on what the viewer actually wants to see. It helps developers stop guessing and start fixing the specific parts of the movie pipeline that are broken.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →