← Latest papers
📊 statistics

The declared winner of a clinical prediction model comparison is often decided by the random split: a simulation study

This simulation study demonstrates that at the small effect sizes typical of clinical prediction model comparisons, the declared "winner" is frequently determined by random data splitting rather than true model superiority, leading to inflated performance margins and significant disagreement between analysts using the same data.

Original authors: Ananya Sharma

Published 2026-08-25
📖 5 min read🧠 Deep dive

Original authors: Ananya Sharma

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of medical research, scientists constantly build new tools to predict who will get sick and who will stay healthy. These tools, known as clinical prediction models, are like sophisticated weather forecasts for the human body, using patient data to guess the likelihood of a future event. To decide which tool is the best, researchers typically pit them against each other on a set of test data. They calculate a score that measures how often the tool gets the prediction right, and the one with the higher score is declared the winner. This process is the standard way the field advances, with new studies frequently announcing that a new machine learning algorithm has beaten an older, simpler method. However, the differences between these tools are often incredibly small, sometimes just a tiny fraction of a percentage point. When the gap between two models is that narrow, the question arises: is the declared winner actually better, or did it just get lucky with the specific group of patients it happened to be tested on?

A recent simulation study by Ananya Sharma explores this exact uncertainty. The research does not involve real patients or new medical data; instead, it uses a computer to generate thousands of fake scenarios where two prediction models are compared. The goal was to see how often the model that wins the competition is truly the superior one, and how often the result is simply a fluke caused by the random way the data is divided. In these simulations, the researchers created two models that were almost identical in their true ability to predict outcomes. They then ran thousands of comparisons, changing the size of the test groups and the specific random split of the data, to see how often the "winner" changed or how much the reported advantage was exaggerated.

The findings reveal a startling level of instability in how these comparisons are currently conducted. When the true difference between two models is very small, which is the typical case in real-world medical studies, the declared winner is often just a product of chance. In one scenario where the models were nearly identical in skill, two different researchers analyzing the exact same data but using different random splits would announce different winners nearly half the time. Even when one model was genuinely slightly better, a small test group meant that the "better" model was only identified correctly about two-thirds of the time. This means that in a significant number of published studies, the conclusion that one tool is superior to another might be wrong, simply because the specific patients chosen for the test set happened to favor one model over the other by random luck.

Another major discovery concerns the size of the victory margin. When a model wins, the study found that the reported difference in performance is almost always inflated, making the gap look much larger than it truly is. For instance, if the true difference between two models was tiny, the winning model would often report a difference that was more than double the actual value. This happens because the process of picking the winner naturally selects the model that got the most favorable random noise in that specific test. It is similar to flipping a coin: if you flip it ten times and get six heads, you might claim the coin is biased, but the result is just a random fluctuation. In the context of these medical models, the "winner's curse" means that the reported success is a mix of real skill and a large dose of statistical luck, leading researchers to believe they have found a breakthrough when they have mostly found a random spike.

The study also examined whether the problem lies in how the models are built or how the data is tested. By keeping the models fixed and only changing the test group, and then by retraining the models on every new split, the researchers found that the instability comes almost entirely from the test group itself. This suggests that making the models more complex or stable will not fix the problem; the issue is that the test groups used in many studies are simply too small to reliably distinguish between models that are very close in performance. When the test group is small, the random noise in the data drowns out the tiny signal of the actual difference.

Ultimately, this research suggests that the current habit of declaring a single winner based on one test run is often misleading. The study recommends that researchers stop treating a single comparison as a final verdict. Instead, they should report the range of uncertainty around the difference and, crucially, repeat the test many times with different random splits to see how often each model wins. If a model only wins slightly more than half the time across many repeats, the honest conclusion is that the two models are effectively indistinguishable with the current data. Until the test groups are large enough to overcome this randomness, or until studies adopt these more rigorous reporting methods, the "winners" of clinical prediction competitions may be less a reflection of medical truth and more a reflection of the roll of the dice.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →