Markerless Pose Estimation for Resistance Training Technique Assessment
This paper presents a markerless pose estimation framework using BlazePose to quantitatively assess resistance training techniques, such as squats and deadlifts, from ordinary video by comparing joint-angle trajectories against a reference, demonstrating its potential for accessible biomechanical analysis despite limitations related to camera orientation and visual occlusion.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of physical training, the difference between building strength and sustaining injury often comes down to a single detail: how a body moves. For decades, experts have relied on two main ways to judge this movement. The first is the human eye, where a coach watches an athlete and offers feedback based on experience. The second is the laboratory, where scientists use expensive cameras and tiny reflective stickers placed on a person's skin to measure every joint with extreme precision. While the laboratory method provides undeniable accuracy, it is locked away in research centers, far from the weight rooms where most people train. This leaves a gap between the high-tech science of movement and the everyday reality of lifting weights.
A new study from the University of Bristol seeks to bridge this gap by asking a simple question: can a standard smartphone camera, without any stickers or special equipment, tell us if a person is lifting safely? The researchers focused on three fundamental exercises—the squat, the bench press, and the deadlift—because these movements involve multiple joints and carry a high risk of injury if performed incorrectly. They wanted to know if modern computer vision, which allows machines to "see" and map the human body from a video, could replace the need for a laboratory visit. The goal was not to create a perfect medical tool, but to see if a practical, accessible system could capture the essential shape of a movement and compare it to a known good example.
To test this idea, the team built a digital framework that takes ordinary video footage and turns it into a story of motion. They used a lightweight computer program capable of spotting thirty-three key points on the human body, such as the shoulders, elbows, and knees, in every single frame of a video. Once these points were identified, the system connected them to calculate the angles of the joints as the person moved. For the squat, for instance, the program watched how the knee bent and how the torso leaned forward. The researchers then created a "gold standard" reference by selecting a video of a perfect repetition from a collection of instructional clips. This perfect movement served as the target, a template against which all other attempts could be measured.
The researchers fed hundreds of videos into their system, filtering out those with poor lighting or bodies that were too hidden from view. They focused on a dataset of eighty-nine high-quality recordings. When they analyzed the squat videos, the system worked remarkably well. It successfully tracked the movement in nearly every single frame, capturing the smooth curve of the knee bending and straightening. The computer could see the difference between a deep, controlled squat and one that was shallow or shaky. By comparing the angle of the knee in a test video against the angle in the perfect reference video, the system assigned a score. A score of one hundred meant the movement matched the ideal perfectly, while a lower score indicated a deviation. In one test involving ten repetitions of a squat, the system identified that the first repetition was the most different from the ideal, while the eighth was the closest match, showing that the technology could spot subtle changes in technique even within a single set of exercises.
However, the study also revealed a significant limitation that anyone using such technology must respect: the angle of the camera matters more than almost anything else. The researchers tested this by having a person hold a static squatting position while they walked around them, filming from the side, the back, and the front. When the camera was positioned directly to the side, capturing the movement in a flat, two-dimensional profile, the angle of the knee looked correct. But as soon as the camera moved to the front or the back, the computer's estimate of the knee angle changed dramatically, even though the person's leg had not moved at all. This distortion was made worse when the lifter was inside a metal weight rack, where the bars blocked parts of the body from view. The system struggled to guess where the hidden joints were, leading to confusing and inaccurate data. This finding ruled out the idea that a camera could be placed anywhere; for the technology to work, the video must be taken from a specific side-on view.
The technology also showed that it works better for some exercises than others. While the squat and the deadlift were tracked with high reliability, the bench press proved difficult. In the bench press, the person lies on their back, and the heavy barbell often blocks the view of the arms and shoulders. Because the computer could not see the key joints clearly, it failed to track the movement in a large portion of the videos. This suggests that while the method is powerful for upright movements where the body is fully visible, it is not yet ready for every type of exercise. The researchers concluded that their system is a promising step toward accessible biomechanics, capable of providing useful feedback for the squat and deadlift, but it is not a magic solution that works in every situation.
Ultimately, this work demonstrates that we are moving closer to a future where high-quality movement analysis is available to everyone, not just those with a laboratory budget. The system can successfully extract the story of a lift from a simple video, identifying when a technique is consistent and when it drifts from the ideal. It can tell an athlete that their first few repetitions were shaky and that they found their rhythm later in the set. Yet, the path forward requires discipline. To get the most out of this tool, users must be careful about where they stand with their camera and what they are filming. The technology is not a replacement for a coach, but it is a new kind of mirror, one that can show us the geometry of our own strength, provided we know how to look at it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.