FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models
This paper introduces FitAQA, a comprehensive benchmark comprising 2,219 videos and 5,512 QA instances across 30 exercises with a unified form error taxonomy, designed to evaluate Multimodal Large Language Models' capabilities in fitness Action Quality Assessment and revealing that current models struggle primarily with visual perception.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are watching a friend try to learn a new dance move from a video. You can easily tell if they are doing it right or wrong, but explaining exactly why their knee is bent the wrong way or why their rhythm is off requires a mix of sharp eyes and specific knowledge. This is the heart of Action Quality Assessment (AQA). While computers have gotten really good at simply recognizing what action is happening (like shouting "Jumping Jack!"), they have historically struggled to judge how well it is being done. Recently, a new type of super-smart computer brain called a Multimodal Large Language Model (MLLM) has arrived. These models can see videos and read text, promising to act like a personal coach that doesn't just say "good job" but can actually critique your form. But here's the big question: Are these AI coaches actually ready to help us, or are they just guessing?
This is exactly what the paper FitAQA sets out to find. The researchers, working with sports science experts, built a massive "gym class" for AI to test its coaching skills. They didn't just ask the AI to give a score; they created a detailed checklist of 30 different bodyweight exercises (like push-ups and squats) and broke down "good form" into six specific categories: alignment, symmetry, stability, coordination, tempo, and completeness. They then fed the AI 2,219 videos and over 5,500 questions to answer. The results were a bit of a reality check: while these AI models are impressive, they are currently terrible at spotting the tiny, crucial details that make an exercise safe and effective. They often miss the errors entirely or can't tell when in the video the mistake happened. The study suggests that the main problem isn't that the AI doesn't know the rules of fitness; it's that it's struggling to see the visual evidence clearly enough to apply those rules.
The Great AI Gym Test
Think of the current state of AI video understanding as a student who has read every textbook on physics but has never actually watched a ball bounce. They know the theory, but they can't predict where the ball will land. The authors of this paper realized that to test if AI can really be a fitness coach, they needed a test that wasn't just about guessing a final score. They needed a test that asked: "Did you see the error?" "Do you know it's an error?" and "Exactly when did it happen?"
To do this, they created FitAQA, a benchmark (a standard test) that is like a giant, digital gym floor. Instead of just looking at one type of exercise, they covered 30 different moves, from full-body cardio like jumping jacks to core strength moves like reverse crunches. But the real magic is in how they organized the mistakes. Instead of making a separate list of errors for every single exercise (which would be chaotic), they worked with human experts to create a unified "Error Taxonomy."
Imagine a dictionary of mistakes that applies to almost everything. Whether you are doing a squat or a push-up, the errors fall into six buckets:
- Alignment: Are your joints lined up correctly?
- Symmetry: Is your left side doing the same thing as your right?
- Stability: Are you wobbling or shaking uncontrollably?
- Coordination: Are your arms and legs moving together in harmony?
- Tempo: Is the rhythm too fast or too slow?
- Completeness: Did you finish the full range of motion, or did you cut it short?
They found 38 specific types of recurring errors across these categories. For example, "Forward Head Posture" or "Insufficient Range of Motion." This allowed them to test the AI's ability to spot these specific patterns across different exercises, rather than just memorizing that "a squat with bad knees is bad."
The Three Levels of the Test
The researchers didn't just ask the AI to give a final grade. They broke the task down into three levels, like a video game with increasing difficulty:
- Perception (The Eyes): The AI is shown a video and asked, "What do you see happening with the hips?" It has to choose between neutral descriptions like "The hips stay below the line" or "The hips rise above the line." This tests if the AI can actually see the body parts moving correctly.
- Judgement (The Brain): Once the AI "sees" the movement, it is asked, "Is this hip movement correct?" This requires combining what it saw with its knowledge of fitness rules.
- Temporal Grounding (The Stopwatch): This is the hardest level. The AI is asked, "Exactly when in this 30-second video did the person fail to lift their hips high enough?" It has to point to the specific seconds where the error occurred.
The Results: AI is a Bad Coach (For Now)
When they ran the test with the most advanced AI models available (including big names like GPT-5.5, Gemini, and various open-source models), the results were surprising. The AI models generally performed only slightly better than a random guesser.
- The Eyes are Weak: The models struggled significantly with the Perception task. Even the best models only got about 54% of the visual details right. This means the AI often couldn't tell if a knee was bent too far or if a back was straight.
- The Brain is Confused: Because the AI couldn't see the details well, its Judgement was also poor. It tended to be overly optimistic, often saying an exercise was "correct" even when it wasn't.
- The Stopwatch is Broken: For the Temporal Grounding task, the models were terrible at pinpointing the exact moment an error happened. They would often guess the entire video was the error, or miss the error completely.
The most interesting discovery came from a "control experiment." The researchers asked: What if we just tell the AI exactly what it saw, and then ask it to judge?
When they gave the AI the correct visual description (the "Perception" answer) and then asked it to make the "Judgement," the AI's performance skyrocketed. For example, one model's ability to spot errors jumped from a low score to over 94% accuracy. This suggests that the AI actually knows the rules of fitness and can reason well, but it is failing because it can't see the visual evidence clearly enough to apply those rules. It's like a brilliant coach who is wearing sunglasses that are too dark to see the athlete's form.
The Bottom Line
The paper concludes that while Multimodal Large Language Models are powerful tools for understanding general video content, they are not yet ready to be reliable fitness coaches. They struggle to detect the subtle, quality-relevant details of human movement and cannot precisely locate when mistakes happen. The authors suggest that before we can trust an AI to correct our workout form, we need to solve the problem of visual perception—teaching the AI to see the tiny details of posture and motion as clearly as a human expert does. Until then, it's probably best to keep a real human trainer in the loop.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.