Do Pathology Foundation Models Learn Biology? Uncertainty-Based Evaluation on AML Cytomorphology
This study demonstrates that while general-purpose vision models may outperform pathology-specific foundation models on standard clustering metrics, the latter's ability to capture biologically meaningful structure is better revealed through stable uncertainty patterns that distinguish between cell maturity states in acute myeloid leukemia.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
In the world of medicine, the microscope has long been the primary tool for understanding the invisible machinery of the human body. When a doctor suspects a blood disorder, they look at a slide of cells, searching for subtle differences in shape, color, and texture that reveal whether a cell is healthy or sick. This task is demanding and prone to human error; even experts can disagree on what they are seeing, and the process is slow. To help, scientists have turned to artificial intelligence, teaching computers to recognize these patterns. Recently, a new generation of powerful computer programs, known as foundation models, has emerged. These are like vast libraries of visual knowledge, trained on millions of images to understand the world. Some are trained on general photographs of nature, while others are trained specifically on medical slides. The big question for doctors and researchers is whether these general-purpose programs, which are incredibly smart about the world at large, can truly understand the specific, complex biology of human cells, or if they need to be trained on medical images to do so.
A team of researchers set out to answer this by testing three different computer models on a specific type of blood cancer called acute myeloid leukemia. This disease occurs when blood cells fail to mature properly, getting stuck in an immature state. The researchers used a dataset containing nearly 17,500 images of individual cells, which had already been sorted by experts into fifteen distinct categories based on their biological development. The goal was to see if the computer models could group these cells together correctly without being told the answers beforehand. They tested a standard model trained on general images, a large model trained on general images but using a more advanced learning method, and a specialized model trained on millions of medical pathology slides.
The results revealed a surprising twist. When the researchers looked at how neatly the computer groups were formed, the general-purpose model trained on natural photographs performed the best. It created tight, well-separated clusters that looked perfect on paper. However, when the researchers checked if these groups actually matched the real biological types of cells, this same model failed completely. It grouped together cells that looked similar on the surface but belonged to entirely different families in the body. In contrast, the model trained specifically on medical slides produced groups that were messier and less geometrically perfect, yet these groups aligned much better with the actual biology of the disease. The specialized model understood the subtle differences that matter to a doctor, while the general model was fooled by superficial similarities.
To get to the heart of why this happened, the researchers developed a new way to test the models: they measured how uncertain the computer was about each cell. In biology, some cells are clearly defined and stable, while others are in a transitional state, changing from one stage to another. These transitional cells are naturally ambiguous. The researchers found that the medical-trained model showed high uncertainty exactly where it should: when looking at these transitional, changing cells. It recognized that these cells were hard to pin down. The general-purpose model, however, showed no such pattern; it was equally confident about everything, failing to sense the biological complexity. This uncertainty turned out to be a more reliable sign of true understanding than the neatness of the groups.
The study also highlighted a critical lesson for how we evaluate these tools. If a researcher only looks at standard scores that measure how tidy the groups are, they would choose the wrong model for the job. The model that looks the best on a standard test is actually the worst at understanding the biology. Furthermore, the researchers found that the performance of the medical model could vary depending on how the experiment was started, but its ability to sense uncertainty remained steady and reliable. This suggests that looking at how a model reacts to ambiguity is a better way to judge its intelligence than simply checking if it sorts things into neat piles. Ultimately, the research shows that training a computer on millions of medical images leaves a distinct mark on how it sees the world, allowing it to grasp the biological reality of disease in a way that general knowledge alone cannot achieve.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.