← Latest papers
💻 bioinformatics

Trustworthy ML/AI for Aging Clocks: Preventing Systematic Prediction Bias in Biological Age Estimation

This paper identifies and addresses the critical issue of systematic prediction bias in machine learning-based aging clocks, which can distort downstream association analyses, by proposing a constrained optimization framework to ensure valid biological age estimation and inference.

Original authors: Lee, H., Ye, Z., Yang, Y., Pan, Y., Maron, B., Wang, Z., Kochunov, P., Thompson, P., Hong, L. E., MA, T., Chen, C., Chen, S.

Published 2026-09-06
📖 4 min read☕ Coffee break read

Original authors: Lee, H., Ye, Z., Yang, Y., Pan, Y., Maron, B., Wang, Z., Kochunov, P., Thompson, P., Hong, L. E., MA, T., Chen, C., Chen, S.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine trying to measure how fast a car is aging by looking at its engine, its paint, and the wear on its tires, rather than simply checking the odometer. This is the essence of biological age research. While we all know our chronological age—the number of years since we were born—our bodies do not always age at the same pace. Some people remain physically robust well into their seventies, while others show signs of decline much earlier. Scientists have developed tools called "aging clocks" to estimate this biological age. These tools analyze complex data from our DNA, blood, or brain scans to predict how old we are on the inside. The goal is to identify who is aging too fast and why, so doctors can intervene to prevent disease. However, a new study suggests that the very computers used to build these clocks might be tricking us, leading to conclusions that are not just wrong, but sometimes the exact opposite of the truth.

The problem lies in how these computer programs, known as machine learning models, are trained. To teach a computer to guess biological age, researchers feed it data from thousands of people, using their known chronological age as the answer key. The computer learns to spot patterns in the biological data that match the years on a calendar. But these programs have a hidden flaw: they are naturally biased toward the average. When a model is unsure, it tends to guess the middle ground. For a young person, this means the computer guesses they are older than they really are. For an elderly person, it guesses they are younger. This is not a random mistake; it is a systematic error that shrinks the extremes toward the center. In the world of aging research, this creates a distorted picture where the oldest people appear to be biologically younger than they are, and the youngest appear older.

Researchers Hwiyoung Lee and Shuo Chen, along with their team at the University of Maryland and other institutions, discovered that this bias does more than just miscalculate a number. It fundamentally breaks the science that follows. When scientists use these flawed clocks to study how lifestyle factors like smoking or diet affect aging, the bias can flip the results entirely. In one real-world test involving brain scans, the standard computer models suggested that people with better cognitive skills had "older" brains, implying that being smart made you age faster. This is a nonsensical conclusion. When the researchers corrected the computer's bias, the result flipped to the expected reality: people with better cognitive skills had "younger," more resilient brains. Similarly, in a study of kidney function, the flawed models suggested that worse kidney health was linked to a younger biological age, while the corrected models showed the logical link: poor kidney health is associated with accelerated aging.

To fix this, the team developed a new method called Unbiased Machine Learning Regression. Instead of letting the computer find the path of least resistance, which leads to that average-shrinking error, they added strict rules to the training process. They forced the model to be perfectly accurate for two specific groups of people: those who are younger and those who are older. By anchoring the model to these two points, the computer is compelled to draw a straight, accurate line through the data, ensuring that a prediction of "older" truly means older, and "younger" truly means younger. When they applied this new method to data from the UK Biobank and the Framingham Heart Study, the systematic errors vanished. The new clocks no longer showed the strange, reversed relationships between health and age.

The findings, published in a preprint, highlight a critical need for caution in how we use artificial intelligence to understand human health. The study demonstrates that simply having a computer model that looks accurate on paper is not enough; if the model is biased, it can lead to dangerous misunderstandings about what causes disease and how to treat it. The researchers have made their new, unbiased tools available to other scientists, offering a way to ensure that the next generation of aging clocks tells the truth about our bodies, rather than a story distorted by the limitations of the software used to build them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →