Estimating VO2max in patients with obesity using machine learning: A retrospective observational cohort study
This retrospective study demonstrates that machine learning models, particularly Bayesian Ridge Regression and Automatic Relevance Determination, can effectively estimate VO2max in adults with obesity using accessible physiological inputs, offering a promising alternative to traditional testing despite a remaining accuracy gap compared to direct measurements.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
The human body is a remarkable engine, but its most vital gauge is often the hardest to read. This gauge is cardiorespiratory fitness, the measure of how efficiently the heart, lungs, and muscles work together to consume oxygen during intense effort. In medical terms, this peak capacity is known as VO2max. It is a powerful predictor of long-term health; a low score carries a risk of early death comparable to smoking or diabetes. For decades, the only way to measure this accurately was to put a person on a treadmill or stationary bike, have them exercise until they could go no further, and analyze their breath with expensive, complex machinery. This process is time-consuming, requires specialized staff, and can be physically daunting, especially for individuals carrying extra weight. Consequently, many people, particularly those with severe obesity, are never tested, leaving doctors without a crucial piece of health information.
A team of researchers in Norway set out to solve this problem by asking if a computer could learn to guess a person's fitness level using only the basic information already available in a standard medical record. They focused specifically on adults with severe obesity, a group often excluded from the studies that created the old, standard formulas for estimating fitness. By feeding data from 253 patients into advanced computer programs, the team trained these models to find patterns between simple measurements—like age, height, and weight—and the actual oxygen consumption measured during rigorous exercise tests. Their goal was not to replace the gold-standard medical test, but to create a reliable, accessible alternative for situations where the full test is impossible to perform.
The researchers approached the problem in three steps, gradually giving the computer more information to work with. In the first step, the model only saw four basic facts: the patient's sex, age, height, and body weight. In the second step, they added measurements of waist and hip size, along with blood pressure readings. In the final, most detailed step, the model also learned about the patient's fat mass and how fast their heart beat at rest and at its maximum. The team tested eleven different types of computer learning models to see which one could make the most accurate guesses. They found that two specific types of models, known as Bayesian Ridge Regression and Automatic Relevance Determination, performed the best. These models learned to weigh the importance of each piece of information differently depending on how much data was available.
When the models were tested on patients they had never seen before, they showed a clear pattern of improvement as more data was added. The simplest model, using only the four basic facts, made an average error of about 12.5 percent when guessing the oxygen consumption. When the model was given the extra details about body shape and blood pressure, the error dropped to 12.0 percent. With the full set of information, including heart rate and fat mass, the error fell further to 11.3 percent. While this is not as precise as the direct measurement from a treadmill test, which usually has an error of only 2 to 6 percent, the researchers noted that an 11 percent margin of error is often acceptable for clinical estimates when the direct test cannot be done. The study also revealed that the computer models were remarkably consistent in what they found important. Across all three levels of complexity, the patient's sex and age were the strongest predictors of fitness. Men consistently showed higher fitness levels than women, and fitness declined as age increased, a pattern that aligns with decades of established medical knowledge.
The study also uncovered how the body's composition influences these estimates. In the simplest model, body weight did not seem to matter much. However, once the model learned about fat mass and other physiological details, body weight became a significant factor. This makes sense because a heavier person needs to move more mass, which requires more oxygen, but only if that extra mass is active tissue rather than just fat. The researchers found that waist circumference was also a useful clue; a larger waist was linked to lower fitness, likely because excess abdominal fat can restrict the movement of the diaphragm and make breathing harder during exercise. Surprisingly, hip circumference and blood pressure readings had very little influence on the final estimates, suggesting that for this specific group of patients, these measurements did not add much new information beyond what was already known from age, sex, and weight.
One of the most promising features of the computer models was their ability to express doubt. Because the best-performing models were built on a type of mathematics that handles probability, they could tell the researchers how confident they were in each guess. The study found a direct link between the size of the error and the model's stated uncertainty: when the model was less sure, it was more likely to be wrong. This feature could be vital in a real-world clinic, as it would allow a doctor to see not just a number, but also a warning sign if that number was likely to be inaccurate. This would help clinicians decide when it is safe to rely on the estimate and when a patient truly needs the full, rigorous exercise test.
The researchers were careful to note that their findings are based on a relatively small group of patients from a single hospital in Norway, and that the models still fall short of the accuracy of direct measurement. They emphasized that while these tools are a step forward, they are not a finished solution. The study suggests that with larger groups of people and perhaps more data on lifestyle habits, these computer guesses could become even sharper. For now, the work demonstrates that machine learning can successfully translate routine medical data into a meaningful estimate of heart and lung health, offering a new, less burdensome way to monitor fitness in a population that has historically been difficult to assess.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.