← Latest papers
🧬 biology

Towards Comprehensive Cellular Characterisation of H&E slides

The paper introduces HistoPLUS, a highly efficient and generalizable deep learning model trained on a novel pan-cancer dataset that significantly outperforms existing state-of-the-art methods in detecting and classifying diverse cell types within tumor microenvironments, including understudied and rare cells, while offering improved cross-domain transferability.

Original authors: Benjamin Adjadj, Pierre-Antoine Bannier, Guillaume Horent, Sebastien Mandela, Aurore Lyon, Kathryn Schutte, Ulysse Marteau, Valentin Gaury, Laura Dumont, Thomas Mathieu, MOSAIC consortium, Reda Belbah
Published 2026-07-27
📖 4 min read☕ Coffee break read

Original authors: Benjamin Adjadj, Pierre-Antoine Bannier, Guillaume Horent, Sebastien Mandela, Aurore Lyon, Kathryn Schutte, Ulysse Marteau, Valentin Gaury, Laura Dumont, Thomas Mathieu, MOSAIC consortium, Reda Belbahri, Benoît Schmauch, Eric Durand, Katharina Von Loga, Lucie Gillet

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are a detective trying to solve a crime, but instead of a messy room, your crime scene is a tiny slice of human tissue. In the world of medicine, doctors use a special stain called H&E (think of it as a high-contrast black-and-white filter) to turn these tissue slices into colorful maps. On these maps, different cells look like different shapes and shades. To understand what's going on inside a patient—like whether a tumor is fighting back or giving up—doctors need to count and identify every single cell in that tiny slice. It's like trying to find specific types of ants in a massive, crowded anthill just by looking at a blurry photo.

For a long time, computers have been terrible at this. They are great at spotting "something that looks like a cell," but they struggle to tell the difference between a soldier cell (immune cell), a construction worker cell (stromal cell), or a bad guy cell (cancer cell), especially when the cells are squished together or look very similar. Existing computer programs are like detectives who only know how to spot the most common ants; if they see a rare or weird-looking ant, they get confused or just guess. This is a big problem because the rare ants often hold the most important clues about how a disease will behave. Scientists have been trying to build better "super-detectives" (AI models) to read these maps, but they've been stuck because they didn't have enough training data with clear labels for all the different types of cells, especially the rare ones.

This is where a new team of researchers steps in with a fresh approach. They built a brand-new, super-smart AI model called HistoPLUS. Think of HistoPLUS as a detective who didn't just learn from a few textbooks but spent years studying a massive, carefully curated library of 108,722 individual cell "snapshots" from six different types of cancer. Unlike previous models that were trained on generic data, this team used a clever "active learning" strategy. Imagine a teacher who notices a student is bad at identifying red beetles, so instead of showing them 100 blue butterflies, the teacher specifically finds and shows them 50 red beetles until they master it. The team did this with rare cell types, ensuring their AI learned to spot the tricky, understudied cells that others missed.

The results are impressive. When they tested HistoPLUS on brand-new, unseen tissue samples, it didn't just do a little better; it crushed the competition. It found cells 5.2% more accurately and classified them 23.7% better than the current best models, all while being five times smaller and lighter (using 5x fewer computer "brain cells" or parameters). This means it's faster and cheaper to run. Most excitingly, HistoPLUS can now identify seven types of cells that were previously too difficult for computers to handle, including specific immune cells and structural cells. Even better, the model showed it could handle two types of cancer it had never seen before during its training, proving it's not just memorizing answers but actually learning the rules of the game.

The researchers argue that while massive, complex AI models (the "giants" of the field) are popular, they aren't always the best solution. Their tests showed that these giant models didn't perform much better than HistoPLUS but required way more computing power. HistoPLUS proves that you don't need a supercomputer to get super results; you just need the right training data and a smart, efficient design. By releasing their model and the code for free, the team hopes to help other scientists unlock new secrets hidden in these tissue maps, potentially leading to better ways to predict how patients will respond to treatments. They aren't claiming to have solved cancer, but they have handed the medical community a much sharper pair of glasses to see the microscopic world more clearly than ever before.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →