CytoFormer: A Molecularly Supervised Cell Foundation Model for Histopathology Cell Classification
CytoFormer is a novel cell foundation model that leverages paired H&E histology and spatial transcriptomics data from 16 organs to achieve accurate, label-efficient, and generalizable cell classification without relying on manual pathologist annotations.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of medicine, the most common way to look inside the human body is through a microscope, examining thin slices of tissue stained with pink and purple dyes. This is the standard of care for diagnosing diseases like cancer. A pathologist looks at these slides and identifies the different types of cells present—whether they are healthy tissue, invading immune cells, or rogue tumor cells. This visual inspection is powerful, but it is also slow, subjective, and difficult to scale. Counting millions of individual cells by eye is impossible, and even experts can disagree on what a specific cell is, especially when different cell types look nearly identical under a microscope. For decades, the dream has been to teach computers to do this counting automatically, but building such a computer has been stuck in a bottleneck: it requires a human expert to label millions of cells by hand to teach the machine what to look for. This manual work is tedious, expensive, and often unreliable, leaving the field without enough high-quality data to train truly intelligent systems.
A team of researchers at the University of Pennsylvania has found a way to bypass this human bottleneck entirely. Instead of asking pathologists to label cells, they used a different kind of biological map to teach the computer. They paired the standard pink-and-purple tissue slides with a newer technology called spatial transcriptomics, which acts like a molecular barcode scanner. This scanner reads the genetic activity inside each individual cell while it is still sitting in the tissue. Because the genetic code is unique to a cell's identity, it tells the researchers exactly what type of cell they are looking at, without any human guessing. By matching these molecular identities to the corresponding images on the standard slides, the team created a massive, self-supervised dataset. They used this to train a new artificial intelligence model called CytoFormer, which learned to recognize cell types just by looking at the shape and texture of the cells in the routine microscope images.
The result is a system that can identify cell types with a level of accuracy that rivals human experts, but without needing a human to draw a single box around a cell during training. The researchers assembled data from 81 different tissue sections covering 16 different organs in the human body, ranging from the breast and lung to the brain and bone marrow. In total, they processed more than 15 million individual cells. For each cell, the computer learned to associate the visual appearance in the stained image with the specific molecular identity provided by the genetic scanner. This allowed the model to learn a universal language of cell shapes that works across the entire body. When tested on new tissue sections that the model had never seen before, it correctly identified the cell types in 85 percent of cases on average. More importantly, the model did not just count cells; it reconstructed the entire architecture of the tissue, correctly mapping out where tumors, immune cells, and healthy structures were located, effectively reproducing the biological landscape of the organ.
What makes this discovery particularly significant is how well the model transfers its knowledge to new situations. The researchers tested whether the AI could handle organs and cell types it had never encountered during its training. They applied the model to four different public datasets containing cells from organs like the colon and skin, which were not part of the original training set. Even without seeing these specific tissues before, the model outperformed six other leading artificial intelligence systems that had been trained on millions of manually labeled images. It was especially effective in difficult scenarios where cells look very similar, such as distinguishing between normal healthy tissue and cancerous tissue that mimics it. In one test, the model learned to spot normal cells hidden among look-alike tumor cells using only a few hundred examples provided by a human, whereas other systems required many more examples to reach the same level of accuracy. This suggests that the model learned the fundamental visual rules of how cells look in their environment, rather than just memorizing a list of specific cell types.
The researchers also explored how much of the surrounding tissue the computer needs to see to make a correct identification. They tested different sizes of the image window around each cell, from a tiny view showing just the cell itself to a large view showing the whole neighborhood. They found that the model worked best when it could see the cell and its immediate neighbors, but not the entire tissue section. This balance allowed the computer to understand the context of the cell—whether it was sitting in a blood vessel, a tumor mass, or a healthy gland—without losing the fine details of the cell's own shape. This specific window size, which corresponds to a very small but clear view of the cell, proved to be the sweet spot for accuracy.
By turning the molecular data from the genetic scanner into a teaching tool for visual recognition, the team has created a reusable resource that can be applied to any routine tissue slide. The model does not require expensive new equipment to run; it can analyze standard slides that are already being used in hospitals every day. The researchers have made the code and the trained model available to the scientific community, allowing others to use this new way of seeing cells. While the current model groups some very similar cells together, such as different types of immune cells, because they look alike, it represents a major step forward. It proves that we can build powerful tools for cellular analysis without relying on the slow and expensive process of manual labeling. This approach opens the door to analyzing millions of cells across thousands of patients, turning routine microscope slides into detailed maps of cellular activity that could help doctors diagnose diseases earlier and more accurately.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.