Can Vision Language Models Judge Action Quality? An Empirical Evaluation
This paper presents a comprehensive empirical evaluation revealing that state-of-the-art Vision Language Models currently perform only marginally above random chance in Action Quality Assessment due to systematic biases and fundamental limitations in fine-grained movement analysis, indicating that significant advancements are required before reliable real-world deployment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have hired a very smart, well-read robot to be a judge for a gymnastics competition, a diving meet, and a gym workout session. You tell this robot, "Look at these videos, and tell me if the athlete is doing the move correctly, or give them a score from 1 to 10."
This paper is essentially a report card on how well these "Vision Language Models" (VLMs)—the smartest AI robots we have right now—actually perform this job. The short answer? They are currently terrible at it.
Here is the breakdown of the study using simple analogies:
1. The Setup: The "Book Smarts" vs. "Street Smarts" Problem
Think of these AI models as super-obsessed students who have read every textbook on gymnastics but have never actually watched a live game.
- The Expectation: Because these AIs can chat, write poems, and understand complex instructions, researchers hoped they could naturally understand human movement. They thought, "If it knows the rules of a push-up from a book, it can spot a bad push-up in a video."
- The Reality: When tested on real videos of people doing squats, diving, or figure skating, the AIs performed barely better than if they had just closed their eyes and guessed.
- The Analogy: It's like asking a chef who has read every cookbook in the world to taste a soup and tell you if it's salty. They might know the theory of salt, but without actually tasting the specific bowl in front of them, they can't judge it.
2. The Experiments: Trying to "Cheat" the System
The researchers tried to help the AIs by giving them different tools, hoping to boost their performance. Here is what they tried and why it failed:
- The "Skeleton" Glasses: They tried showing the AI just the stick-figure skeleton of the person instead of the full video, hoping it would focus on the bones and angles.
- Result: It was like giving a detective a map of a crime scene but hiding the actual bodies. The AI got confused because it couldn't see the muscle movement or the "flow" of the action.
- The "Prompt" Hacks: They tried changing the questions.
- The "Grounding" Trick: They told the AI, "Look at the video first, then answer."
- The "Step-by-Step" Trick: They told the AI, "First describe what you see, then think, then answer."
- The "Example" Trick: They showed the AI a few examples of good and bad moves before asking it to judge a new one.
- Result: These tricks helped a tiny bit in very specific situations (like looking at a single photo), but for videos, the AI still stumbled. It's like telling a student, "Read the question twice before answering," which helps a little, but doesn't fix the fact that they don't understand the subject matter.
3. The Two Big Flaws (The "Biases")
The study found that the AIs have two major personality quirks that ruin their judging:
Flaw A: The "Optimist" Bias
The AIs have a strong tendency to assume everything is being done correctly, even when it's clearly wrong.
- The Analogy: Imagine a parent watching their child ride a bike. Even if the child is wobbling and about to crash, the parent says, "Great job! Perfect balance!" The AI does this because it "knows" from its training data that people usually try to do exercises correctly, so it defaults to "Correct" unless forced to think otherwise.
Flaw B: The "Word-Sensitive" Bias
The AIs are easily tricked by how a question is phrased.
- The Analogy: If you ask, "Is the person's back straight?" the AI might say "Yes." But if you ask, "Is the person's back not straight?" the AI might suddenly say "No," even if the video is exactly the same. It's reacting to the words in the question rather than the video in front of it. It's like a student who answers "True" just because the sentence sounds positive, without actually checking the facts.
4. The "Head-to-Head" Test
To see if the AIs could at least tell the difference between a "good" move and a "bad" move, the researchers showed them two videos side-by-side and asked, "Which one is better?"
- The Result: The AIs got better at this only when the difference was huge (like comparing a professional diver to someone falling off a diving board). But when the difference was subtle (like a professional diver making a tiny splash vs. a perfect entry), the AIs couldn't tell the difference. They were essentially guessing.
The Conclusion: Don't Hire the Robot Yet
The paper concludes that while these AI models are amazing at writing stories and answering trivia, they are fundamentally broken when it comes to judging fine details of human movement.
- Why? They rely too much on what they think should happen (their "book smarts") and not enough on what they are actually seeing (their "street smarts").
- The Takeaway: We cannot trust these models to replace human coaches or judges in physical therapy or sports yet. Before we can use them in the real world, we need to teach them how to actually watch and analyze movement, rather than just guessing based on the words they read.
In short: The AI is a brilliant librarian who knows every rule of the game, but it's currently a terrible referee because it can't actually see the players on the field.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.