Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models
This paper addresses the lack of robust benchmarks and high-quality data in video understanding reward modeling by introducing the Video Understanding Reward Bench (VURB), the large-scale Video Understanding Preference Dataset (VUP-35K), and state-of-the-art discriminative and generative reward models (VideoDRM and VideoGRM) that significantly advance performance and reasoning capabilities.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are the head coach of a team of AI robots. These robots are learning to watch videos and answer questions about them. Sometimes, they get the answer right but explain it poorly; other times, they get the answer wrong but sound very confident.
Your job is to be the Referee. You need to watch the robots' answers and decide which one is better so the team can learn from the mistakes. This paper is about building a better Referee specifically for video tasks.
Here is the story of what the authors built, explained simply:
1. The Problem: The Referees Were Blind
The authors noticed that while AI has gotten really good at judging text (like essays) and images (like photos), it is terrible at judging videos.
- The Old Tests Were Flawed: Existing tests were like a pop quiz with only 5 questions. They didn't cover enough topics, and the questions were too short. It was like trying to judge a marathon runner by watching them tie their shoes.
- The Data Was Missing: To train a good referee, you need thousands of examples of "Good Answer" vs. "Bad Answer." For videos, this data was almost non-existent because it's hard and expensive to make.
- The Result: The current AI referees were guessing. They were barely doing better than flipping a coin (50% accuracy). They couldn't tell the difference between a robot that actually understood the video and one that just guessed the right letter.
2. The Solution: Building a Better Stadium and a New Rulebook
To fix this, the team built a complete new system with three main parts:
A. The New Stadium: VURB (The Benchmark)
They built a massive, high-quality testing ground called VURB.
- The Scale: Instead of 5 questions, they created 2,100 complex video challenges.
- The Depth: They didn't just ask "What happened?" They asked the robots to explain why it happened, step-by-step. Imagine asking a student to not just solve a math problem, but to write out their entire thought process. The answers in this test are very long (over 1,000 words on average) to force the AI to really think.
- The Fairness: To stop the AI from cheating by just picking the first answer it sees, they used a "Majority Vote" rule. They ask the AI to judge the same video 8 times, swapping the order of the answers each time. If the AI picks the right one most of the time, it passes.
B. The New Training Camp: VUP-35K (The Dataset)
You can't train a referee without practice games. Since no one had enough practice data, they built a machine to create it.
- The Factory: They created a fully automated pipeline (no humans needed) that generated 35,000 pairs of "Good Answer" vs. "Bad Answer" for videos.
- The Trick: To make sure the "Bad Answers" were actually bad, they sometimes intentionally gave the video a little "glitch" (like lowering the quality or removing frames) to trick the AI into making mistakes. This created a huge library of examples showing exactly how and why AI fails at video reasoning.
C. The New Referees: VideoDRM and VideoGRM
Using their new training data, they built two new types of referees:
- VideoDRM (The Scorekeeper): This referee looks at two answers and gives them a score (like 9.8 vs. 2.9) to decide which is better.
- VideoGRM (The Commentator): This referee doesn't just pick a winner; it writes a detailed explanation of why one answer is better, acting like a sports analyst breaking down a play.
3. The Results: The New Referees Win
When they put their new referees to the test:
- Old Referees: The best existing AI models scored around 55-60%. They were barely better than random guessing.
- New Referees: Their new models (VideoDRM and VideoGRM) jumped to 63-64%.
- The Comparison: Their new models beat even the most expensive, "closed-source" commercial AI models (like the latest versions of GPT or Gemini) on these specific video tasks.
4. The "Best-of-N" Superpower
The paper also showed that these new referees help the AI team get smarter during the actual game.
- Imagine the AI robot is asked a hard question. Instead of just giving one answer, it generates 8 different possible answers.
- The new referee looks at all 8, picks the best one, and the robot gives that one as the final answer.
- The Result: This simple trick made the AI significantly more accurate, proving that having a good referee is just as important as having a smart player.
Summary
In short, the authors said: "We can't judge video AI well because we don't have good tests or enough practice data." So, they built a giant new test, a factory to make practice data, and two new AI referees that are now the best in the world at watching videos and deciding which answers make sense.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.