Identical runs disagree more than different models: reproducibility limits in deep learning for brain-tumour MRI classification
This study demonstrates that uncontrolled stochastic variations in nominally identical deep learning runs for brain-tumour MRI classification can exceed the performance differences between distinct model architectures, thereby undermining the reliability of standard ablation studies and necessitating stricter reproducibility protocols.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the high-stakes world of medical imaging, computers are increasingly asked to look at pictures of the human brain and spot tumors. These pictures, called MRI scans, are complex and detailed, and the software used to analyze them relies on a type of artificial intelligence known as deep learning. To train these systems, researchers feed them thousands of images and let the computer learn patterns on its own, adjusting its internal settings millions of times until it gets better at identifying disease. The goal is to create a tool that doctors can trust to help diagnose conditions like brain tumors. However, for this technology to be safe and useful, the results must be consistent. If a computer is shown the same data twice, it should ideally give the same answer. When scientists compare two different computer models to see which one is better, they often look for very small differences in accuracy, sometimes just a fraction of a percent, to decide which model should move forward to real-world testing.
A researcher at San Francisco Bay University set out to test a specific idea: whether adding a special type of reasoning to a standard brain-imaging computer model would make it better at spotting tumors. The standard model looks at the image directly, while the new version tried to understand the relationships between different parts of the image, similar to how a human might connect the dots between features. The researcher built a rigorous experiment to test this, running the computer models over and over again to see if the new approach truly offered an advantage. What he found, however, was not a breakthrough in tumor detection, but a startling discovery about the reliability of the testing process itself. He discovered that the computer's own training process was so unstable that two identical runs of the exact same model could produce different results that were larger than the tiny differences he was trying to measure between the two different models.
The study began with a large collection of MRI slices from hundreds of patients, divided carefully so that no single patient's data appeared in both the training and testing groups. This was crucial because if the computer saw the same patient's tumor in both groups, it would simply memorize that specific person rather than learning to recognize the disease generally. The researcher trained a standard model and a more complex version that included the extra reasoning layer. He ran each version multiple times, changing a random starting number, known as a seed, to see how the results varied. In the world of computer science, this seed is supposed to ensure that if you start with the same number, the computer does the exact same thing every time. The expectation was that the complex model would either be slightly better or slightly worse than the simple one, and that running the test three times would give a clear average.
Instead, the results showed a chaotic pattern. When the researcher ran the exact same model configuration twice with the same starting number, the accuracy scores differed by more than two percentage points on average. This variation was so large that it completely drowned out the small differences between the two different models he was comparing. In fact, the difference between the two identical runs was more than twice as big as the difference between the simple model and the complex one. This meant that the experiment was essentially measuring noise rather than signal. The researcher realized that the computer was not behaving as a reliable instrument; it was as if a scale used to weigh gold coins gave a different reading every time you stepped on it, even if you stood on the exact same spot.
Digging deeper, the researcher found two main reasons for this unreliability. First, the computer was being set up incorrectly. The random starting number was being applied after the model had already been built, meaning the model's initial settings were still random and uncontrolled. Even after fixing this mistake and applying the starting number before the model was built, the results were still not perfectly identical. The second issue was hidden deep inside the computer's mathematical engine. Even though the software claimed to be running in a perfectly deterministic mode, a specific part of the calculation used to teach the model was producing slightly different results every time. This hidden flaw was invisible to the standard checks that researchers use to ensure their experiments are reproducible.
The study also revealed that the way data is split can dramatically change the outcome. When the researcher split the data by individual MRI slices rather than by patient, the computer's accuracy jumped by nearly five percentage points. This happened because the computer was essentially recognizing patterns from the same patient in both the training and testing sets, allowing it to recognize the patient rather than the tumor. This inflated score was an illusion, masking the model's true ability. Furthermore, when the researcher tested how well the models handled distorted or noisy images, a single test run suggested one model was more robust, but when the test was repeated, the result flipped, showing the other model was better. This reversal proved that a single experiment is not enough to draw a conclusion.
The final conclusion of the work was that the extra reasoning layer did not improve the ability to detect brain tumors. The complex model was not better, and in some ways, it was slower and less reliable. More importantly, the study highlighted a critical flaw in how many medical AI studies are conducted. The differences researchers often claim to find are smaller than the natural variation in the testing process itself. The researcher argued that before scientists can trust a new model, they must first verify that their testing setup is stable enough to detect small changes. This involves checking that the random starting numbers are applied correctly and running the same test twice to measure the natural variation. These checks cost very little time and effort, yet they are often skipped. Without them, the field risks building medical tools on foundations that are too shaky to support the weight of clinical decisions. The paper serves as a reminder that in the pursuit of better medical technology, ensuring the measurement tool is accurate is just as important as the tool itself.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.