← Latest papers
💻 computer science

Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation

The paper introduces JudgeFit, an iterative framework that constructs personalized evaluation taxonomies for individual vision-language models to better assess physical consistency in video generation, demonstrating a 32% improvement over global schemas while revealing model-specific blind spots.

Original authors: Yu Cao, Ziquan Liu, Zhensong Zhang, Jiankang Deng, Shaogang Gong, Jifei Song

Published 2026-06-23
📖 4 min read☕ Coffee break read

Original authors: Yu Cao, Ziquan Liu, Zhensong Zhang, Jiankang Deng, Shaogang Gong, Jifei Song

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are running a talent show for video generators. You have 16 different judges (AI models) to help you decide which videos are physically realistic and which ones look weird.

The problem is, every judge sees the world differently.

The Old Way: One Size Fits All (The "Bad" Ruler)

Traditionally, researchers gave every judge the exact same checklist to grade the videos. It was like handing a carpenter, a chef, and a musician the same ruler and asking them to measure a cake.

  • The carpenter might only look at the height.
  • The chef might only look at the texture.
  • The musician might not know what to do with the ruler at all.

In the paper, they found that when they used a "global" checklist (a fixed set of rules like "gravity," "collisions," and "lighting"), most AI judges ignored three-quarters of the list. They just focused on one thing (usually "mechanics") and gave the rest a zero. It was like asking a colorblind person to judge a painting based on a list of 10 different colors; they just pick the one they can see and ignore the rest.

The New Solution: JUDGEFIT (The "Custom Tailor")

The authors created a new system called JUDGEFIT. Instead of forcing every judge to use the same ruler, JUDGEFIT builds a custom ruler for each specific judge based on what that judge is actually good at seeing.

Here is how it works, step-by-step:

Step 1: The "Seed" (Asking the Judge to Talk)

First, they show a small set of videos to an AI judge and ask: "What looks wrong here? Just tell me in your own words."

  • If the judge says, "The ball floated like a balloon," the system notes "floating" as a rule.
  • If the judge never mentions "shadows," the system knows this judge is blind to shadows, so it doesn't put "shadows" on their custom checklist.
  • The Analogy: It's like asking a new employee, "What mistakes do you notice?" and building their job description based only on the mistakes they actually spot.

Step 2: The "Refine" (The Coach and the Player)

Now, the system puts the judge to work using this new custom checklist. But here's the trick:

  • The Player (The Video AI): Scores the videos based on the custom checklist. It doesn't know what the "right" answer is; it just follows the rules.
  • The Coach (A separate AI Editor): Looks at the Player's scores and compares them to what humans thought.
    • If the Player says a video is perfect, but humans hated it: The Coach says, "You missed something! Let's add a new rule to your checklist."
    • If the Player is checking two things that mean the same thing: The Coach says, "You're doing double work. Let's merge these rules."
    • If the Player is checking something that doesn't match human opinion: The Coach says, "Stop checking that; it's confusing your score."

This happens in a loop. The Coach tweaks the checklist, the Player tries again, and they keep going until the Player's scores match the humans' opinions as closely as possible.

The Results: Why It Matters

The paper tested this on 16 different AI models (from small open-source ones to massive closed-source ones).

  • The Win: For every single judge, the custom checklist worked better than the standard "one-size-fits-all" checklist. On average, the scores became 32% more accurate at predicting what humans would think.
  • The Surprise: They found that even the "smartest" AI judges have blind spots. One might be amazing at spotting gravity errors but terrible at spotting reflection errors. Another might be the opposite. The old system treated them all as "good" or "bad" overall, but JUDGEFIT revealed their specific strengths and weaknesses.

The Bottom Line

The paper argues that we shouldn't force every AI judge to use the same rules. Just like you wouldn't use a fishing rod to fix a car, you shouldn't force an AI to judge videos using rules it can't perceive.

JUDGEFIT is a method that asks each AI, "What can you see?" and then builds a unique, custom grading system for it. This makes the AI a much better, more reliable judge of physical reality in videos.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →