An ablation measured twice: small component effects can change sign across repeated training runs in breast-cancer histopathology segmentation
This study demonstrates that in breast-cancer histopathology segmentation, the measured contributions of small architectural components in ablation studies can vary significantly in magnitude and even change direction across repeated training runs, highlighting the critical need for replication before drawing definitive conclusions about component efficacy.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the field of medical imaging, computers are increasingly tasked with reading microscope slides to find cancer. These digital tools, known as deep learning models, are trained to recognize patterns in tissue samples that the human eye might miss. To build these tools, researchers combine various techniques, such as adjusting the colors of the images to account for differences between laboratories, or teaching the computer to look at the image in different levels of detail at once. When a new model is built, scientists create a summary table to show which of these techniques actually helped. This table, often called an ablation study, lists the specific contribution of each part, telling the reader, "this piece added value, that piece did nothing." The goal is to separate the useful tools from the noise so that future doctors and researchers know exactly what to trust.
However, a hidden problem lurks beneath these summaries. Even when a computer program is run twice with the exact same settings, the results are rarely identical. Tiny, invisible fluctuations in the computer's calculations can cause the final score to wiggle up or down. If this natural wiggle is large enough, it can make a helpful tool look useless, or a useless tool look helpful, simply by chance. This uncertainty makes it difficult to know if a reported improvement is a real discovery or just a lucky roll of the dice.
A researcher at San Francisco Bay University decided to test how much this natural wiggle matters in the specific task of identifying breast cancer in tissue samples. The study focused on a popular benchmark dataset containing thousands of image patches from patient slides. The researcher built a computer model designed to spot tumors and then tested seven different versions of it. Each version turned on or off three specific features: a method to standardize the colors of the images, a way to look at the image in multiple sizes simultaneously, and a mechanism that lets the computer compare fine details with broad overviews. The goal was to see which features truly improved the model's ability to find cancer.
To get a true sense of the computer's natural instability, the researcher first ran a single version of the model twelve times in a row, changing nothing but the random starting point of the calculation. The difference between these paired runs was measured, revealing that the model's score naturally shifted by a small but consistent amount every time it was restarted. This established a baseline for how much the numbers could move without any real change to the model itself.
With this baseline in hand, the researcher then ran the full experiment of seven different model versions. However, the subsequent repetition of the entire experiment was not originally planned as a design choice; it was necessitated because the per-patch prediction files from the first measurement were lost to a version-control mistake. To recover the data, the researcher repeated the grid, performing twenty-one fresh training runs on the same seven configurations using the same data and starting points. This created a second, independent set of results to compare against the first.
The comparison revealed a striking pattern. The large improvements seen in the first round mostly held up in the second round. For instance, the combination of color standardization and multi-size viewing consistently improved the model's performance, though the exact amount of improvement was slightly smaller the second time. However, the smaller, more subtle effects were completely unreliable. One specific feature, which acts like a bridge between fine details and broad views, showed a negative effect in the first round, suggesting it hurt the model's performance. In the second round, that same feature showed a positive effect, suggesting it helped. Another small effect flipped from positive to negative.
The study found that the size of the shift between the two rounds was roughly the same for every feature, regardless of how big the feature's effect was. This means that when a feature's effect is small, the natural wiggle of the computer is large enough to completely drown it out or even reverse its direction. The researcher concluded that for small effects, a single experiment is not enough to draw a conclusion. If a feature's impact is comparable to the natural noise of the system, its result is not stable enough to be trusted.
The only finding that survived this rigorous test was a specific behavior of the color standardization tool. While it did not change the overall average score of the model very much, it consistently helped the weakest hospitals in the dataset the most, bringing their performance up to match the stronger ones. This pattern was clear in both the first and second runs, proving that the tool was genuinely useful for making the system fairer across different locations, even if the overall numbers didn't look dramatic.
Ultimately, this work serves as a cautionary tale for how we interpret computer science results. It shows that when a computer model's improvement is small, it is impossible to tell if it is a real breakthrough or just a random fluctuation without running the experiment multiple times. The study argues that before declaring a specific part of a model to be good or bad, researchers must first measure how much the numbers wiggle on their own. If the effect is smaller than that wiggle, the result is not yet a fact, but merely a suggestion that needs further proof.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.