← Latest papers
📄 medicine

Alzheimer disease risk estimates lose clinical calibration across study phases despite useful patient ranking

Although risk models for Alzheimer's disease maintain useful patient ranking across different study phases, their absolute probability estimates suffer from significant calibration drift, necessitating revalidation of probability scales when recruitment or measurement practices change.

Original authors: Maurice Antony Ewing

Published 2026-09-04
📖 6 min read🧠 Deep dive

Original authors: Maurice Antony Ewing

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a doctor trying to predict the future for a patient with mild memory loss. The goal is to estimate the chance that this person will develop Alzheimer's disease within the next two years. This prediction serves two distinct purposes. First, it helps sort patients into a list, placing those at highest risk at the top so doctors know who needs the most urgent attention. Second, it provides a specific number, a percentage that represents the actual likelihood of the disease appearing. This number guides critical decisions, such as whether to order expensive brain scans, enroll a patient in a clinical trial, or have a serious conversation about the future. For this system to work, the list must be accurate, but the number must also be true to the specific group of people being treated. If the number is wrong, a patient might be told they have a high chance of illness when they are actually safe, or vice versa, leading to unnecessary worry or missed care.

A new study by Maurice Antony Ewing investigates whether these probability numbers remain reliable when the group of patients changes over time. The research focuses on the Alzheimer's Disease Neuroimaging Initiative, a massive, long-running project that has been tracking thousands of people since 2004. Over the years, the project has gone through different phases, recruiting different types of people and using slightly different methods to measure memory and thinking skills. The researcher asked a simple but vital question: if a computer model is trained to predict risk using data from the early years of the project, will it still give accurate percentage estimates when applied to patients from the later years, even if it can still correctly rank who is sicker than whom?

To find the answer, the researcher built eleven different computer models using data from the first phase of the project. These models learned to look at a patient's age, genetic markers, and scores on memory tests to predict the chance of developing Alzheimer's within two years. Once trained, these models were sent to the second phase of the project to make predictions for a new group of patients. The researcher then did the reverse, training models on the second phase and testing them on the first. In both directions, the models were compared against new models built specifically for the target group, which served as the local standard. The study involved over one thousand participants with mild cognitive impairment, tracking whether they progressed to Alzheimer's disease within the two-year window.

The results revealed a clear and troubling split between ranking and risk estimation. When the models trained on one phase were tested on the other, they remained quite good at sorting patients. They could still identify which individuals were more likely to develop the disease than their peers, often performing slightly better than the models built locally for that specific group. However, the specific numbers they assigned to those risks were wildly off. In the second phase, a model trained on the first phase predicted risks that were, on average, fifteen percentage points away from what actually happened, whereas a model built for that phase was only off by about four percentage points. This gap means that a patient might be told they have a forty percent chance of developing the disease when their actual chance is only twenty-five percent, or the reverse. This error is large enough to push a patient across a medical threshold, potentially sending them down a path of unnecessary testing or denying them a trial they could have joined.

The study went further to understand why this happened. The researcher tested whether the error was caused by missing data, such as when certain memory test scores were not recorded in the later phase, or by the inclusion of a specific group of patients with very mild symptoms. By removing these variables and specific patient groups from the analysis, the researcher found that the problem persisted. Even when the models used only the most complete data and excluded the mildest cases, the numbers remained inaccurate. The error was not a glitch in the math or a flaw in a single test; it was a fundamental shift in what the numbers meant. The population in the later phase was different: they were younger, had less severe symptoms at the start, and were less likely to develop the disease within two years compared to the earlier group. A model trained on the older, sicker group simply applied the old rules to a new reality, causing the probabilities to drift.

To fix this, the researcher tested whether the models could be adjusted after they were moved to the new group. They found that simply correcting the average risk level of the new group helped significantly, bringing the error down from fifteen percentage points to about three or four. However, this was not enough to fully fix the problem. The relationship between a patient's specific symptoms and their risk also changed. Only when the model was adjusted to account for both the average risk of the group and the way individual symptoms mapped to risk did the predictions become accurate again. This two-step correction restored the reliability of the numbers, proving that the models themselves were not broken, but rather that they needed to be recalibrated for the new environment.

The findings offer a clear lesson for the future of medical prediction. A computer model that successfully ranks patients by risk is not automatically ready to give them a specific percentage chance of disease. As medical programs evolve, recruiting different people or using new measurement tools, the meaning of a risk percentage changes. A number that was accurate for a group of patients ten years ago may be misleading for a group today, even if the symptoms look similar. The study concludes that whenever a prediction model is moved to a new setting or a new time, doctors and researchers must re-check the accuracy of the probability numbers. They cannot assume the scale remains the same. While the ability to sort patients from high to low risk may travel well, the specific estimate of danger requires a fresh look to ensure it reflects the reality of the people currently in the clinic.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →