← Latest papers
💻 computer science

When Lifts Predict Lifts: Leakage-Controlled Strength Prediction and a Closed-Form Leakage Diagnostic

This paper exposes how machine learning models for predicting athletes' maximal lifts suffer from significant accuracy inflation due to concurrent-lift data leakage, and proposes a closed-form diagnostic to quantify this leakage severity alongside a cost-effective generative model that provides calibrated, physiologically consistent predictions without relying on sibling lift data.

Original authors: Yasin Alipour, Reza Pourgholi

Published 2026-09-24
📖 6 min read🧠 Deep dive

Original authors: Yasin Alipour, Reza Pourgholi

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of strength sports, from the Olympic platform to the local gym, knowing how much an athlete can lift is the foundation of everything. Coaches use these numbers to design training programs, track progress, and decide if a movement is safe to attempt. Ideally, a coach would measure every single lift an athlete can perform—the clean and jerk, the snatch, the back squat, and the deadlift—to get a complete picture. But measuring everything is slow, exhausting, and often impossible for a new athlete who has never competed before. This has led to a growing interest in using computer programs to predict these lifts based on simple facts like age, sex, height, and body weight. The hope is that a machine could look at a person's basic profile and tell you exactly how much they can lift, saving time and effort.

For years, computer models have claimed to do this with stunning accuracy, often getting the numbers right more than 90 percent of the time. However, a new study suggests that these high scores are misleading. The researchers found that the models were not actually learning how to predict strength from a person's body; instead, they were relying on other known lifts as clues to guess the missing one. If a model knows how much someone can squat, it can guess their deadlift with near-perfect accuracy because the two movements are so closely linked. This is like trying to guess a person's height by asking how tall they are; the answer is easy, but it doesn't tell you anything new. The real challenge, the one coaches actually face, is predicting a full profile for a stranger who has never lifted a weight before, using only their age and size.

Yasin Alipour and Reza Pourgholi, researchers at Damghan University in Iran, set out to expose this flaw and build a better way to predict strength. They started by taking a large, public database of CrossFit athletes and reproducing the high-accuracy results that previous studies had reported. When they allowed their computer models to see an athlete's other lifts while making a prediction, the models performed brilliantly, matching the 93 percent accuracy seen in earlier work. But then they changed the rules. They removed the other lifts from the input, forcing the model to predict a lift using only the athlete's age, sex, height, and weight. The results collapsed. The accuracy dropped by a significant margin, falling from the high 90s down to the mid-60s. This dramatic drop revealed that the previous success was largely an illusion created by data leakage, where the answer was hidden inside the question.

The researchers realized that the honest problem is much harder than the one being solved in the lab. Predicting a lift for a new athlete without any history is a difficult task, and the best computer programs currently available only get it right about 65 percent of the time. This is not a failure of the technology, but a reflection of reality: a person's body size and age can only tell you so much about their strength. The study shows that while machines can screen athletes and give a rough starting point, they cannot replace the actual measurement of a lift. The extra information provided by complex algorithms adds only a small amount of value over a simple calculation based on sex and body weight alone.

To fix the confusion in the field, the team developed a new tool to detect this kind of reliance before it happens. They created a simple score that can look at a dataset and predict how much accuracy will be lost if the "sibling" targets—like the other lifts—are removed. This score works by analyzing how the different lifts relate to each other after accounting for basic body facts. If the score is high, it warns that the dataset is likely inflated by leakage. The researchers tested this diagnostic on strength data and found it worked perfectly. They also tried it on completely different fields, such as predicting the strength of concrete mixtures and the energy usage of buildings. In cases where the targets were truly linked, like the different lifts in a gym, the score correctly predicted a large drop in accuracy. In cases where the targets were already determined by the input data, like the heating and cooling needs of a building, the score correctly predicted almost no drop, proving that the tool can distinguish between real prediction and data leakage.

Beyond just exposing the problem, the researchers built a new type of model designed to handle the honest, difficult task of cold-start prediction. Instead of treating each lift as a separate problem, they built a single system that views strength as a combination of overall capacity and a specific style. Imagine strength as a volume knob that controls how strong an athlete is in general, and a set of sliders that determine how that strength is distributed across different movements. This system can take a few known lifts and fill in the gaps for the unknown ones, or predict a full profile for a new athlete. It does this in a fraction of the time and cost of older, more complex methods, while also providing a measure of uncertainty so coaches know how much to trust the number.

The team validated these findings on a massive scale, testing their methods on nearly 300,000 competitive powerlifters from a different database. The results held up: the high accuracy of the old models vanished when the other lifts were removed, and the new diagnostic tool correctly identified the leakage. They also showed that the relationship between lifts is consistent across different populations. A model trained on recreational CrossFit athletes could successfully predict the relationship between lifts for elite powerlifters, even though the elite athletes were much stronger. This suggests that while the absolute numbers change, the underlying structure of how strength is distributed remains the same.

The study concludes with a clear message for coaches and data scientists: the headline numbers in strength prediction are often too good to be true. When a model uses an athlete's other lifts to guess a new one, it is not demonstrating intelligence; it is simply exploiting a known physical relationship. The true value of machine learning in this field lies in its ability to provide a reasonable starting point for new athletes and to fill in missing data when some lifts are known, all while honestly admitting the limits of its certainty. By separating the easy, leaked predictions from the hard, honest ones, the researchers have provided a clearer path forward for using data to understand human strength.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →