← Latest papers
💬 NLP

RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees

This paper introduces RefereeBench, the first large-scale benchmark comprising 11 sports and over 6,000 QA pairs to evaluate MLLMs as sports referees, revealing that even state-of-the-art models struggle with rule application and temporal grounding, achieving only around 60% accuracy and highlighting the significant gap between current capabilities and reliable automated officiating.

Original authors: Yichen Xu, Yuanhang Liu, Chuhan Wang, Zihan Zhao, jinghan luo, Jianzhe Ma, Wenxuan Wang, Qin Jin

Published 2026-04-20
📖 4 min read☕ Coffee break read

Original authors: Yichen Xu, Yuanhang Liu, Chuhan Wang, Zihan Zhao, jinghan luo, Jianzhe Ma, Wenxuan Wang, Qin Jin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart robot that has watched every movie, TV show, and home video ever made. It can describe what's happening in a video better than almost anyone. But now, you ask it to do something much harder: be a sports referee.

This paper, titled "RefereeBench," is like a giant, high-stakes exam for these AI robots to see if they are ready to blow the whistle on a soccer field, a basketball court, or an ice rink.

Here is the breakdown of what the researchers did, using some simple analogies:

1. The Problem: The "Movie Buff" vs. The "Referee"

Think of current AI models (MLLMs) as super-enthusiastic movie buffs. They are great at saying, "Hey, I see a guy running!" or "Oh, that looks like a fight!"

But being a referee isn't just about seeing things; it's about judging them against a rulebook.

  • Movie Buff: "That guy hit the other guy with his stick."
  • Referee: "That was a 'slashing' foul because he hit him with the stick above the waist during a specific play, which means the team gets a 2-minute penalty."

The researchers wanted to know: Can these AI "movie buffs" actually learn the complex rulebooks of 11 different sports and make fair calls?

2. The Exam: "RefereeBench"

To test this, the team built RefereeBench, which is like a massive, multi-sport final exam.

  • The Classroom: They gathered 925 real video clips from 11 different sports (Soccer, Basketball, Ice Hockey, Tennis, etc.).
  • The Teachers: They didn't just ask random people to grade the videos. They hired certified, real-life referees to watch the clips and write the questions and answers.
  • The Test: The AI had to answer over 6,400 questions. These weren't just "What color is the shirt?" questions. They were tricky ones like:
    • "Was this actually a foul, or just a hard hit?"
    • "What specific rule was broken?"
    • "Exactly when did the bad thing happen?" (This is called "temporal grounding"—like finding the exact second a goal was scored).

3. The Results: The AI is Still a Rookie

The results were a bit of a reality check. Even the smartest AI models in the world (like the ones from Google, OpenAI, and ByteDance) struggled.

  • The Score: The best AI models only got about 60% of the answers right. The best open-source model (Qwen3-VL) got about 47%.
  • The Analogy: If you took a smart high school student and put them in a referee uniform, they might know the rules of the game, but they would still miss the subtle calls that a pro makes. They are "rookies" at best.

Where did they fail?

  1. The "Over-Call" Problem: The AI tends to be too harsh. If a player bumps into another player, the AI often screams "FOUL!" even when it was just normal contact. It's like a nervous referee who blows the whistle for every little touch.
  2. The "Timing" Issue: They are bad at pinpointing the exact moment a foul happened. They might say the foul happened 5 seconds before it actually did.
  3. The "Rulebook" Gap: They can see the action, but they struggle to connect it to the specific, complex rule that makes it illegal.

4. The Surprises

  • Audio Helps: When the AI could hear the video (the crowd cheering, the referee's whistle, the players shouting), it got much better. It's like how a human referee listens for the "crack" of a bat or a whistle to help make a call.
  • Reading the Rules Isn't Enough: The researchers tried giving the AI the actual rulebooks to read while it watched the video (like a cheat sheet). Surprisingly, this didn't help much. The AI still couldn't figure out how to apply the rule to the messy, real-life video. It's like giving a student a dictionary but still having them fail the essay because they don't know how to use the words in a sentence.

5. The Conclusion

The paper concludes that AI is not ready to replace human referees yet.

While these robots are amazing at describing what they see, they aren't ready to make the tough, split-second decisions that determine who wins a championship. They lack the "common sense" and deep understanding of the spirit of the game that human referees have.

The Takeaway:
Think of current AI as a very fast, very knowledgeable intern. It can organize the files and point out the obvious stuff. But until it learns to be a seasoned judge, we still need the human referees to blow the whistle and keep the game fair.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →