← Latest papers
🤖 AI

Mitosis Detection in the Wild: Multi-Tumor and Context-Aware Generalization in the MIDOG 2025 Challenge

The MIDOG 2025 challenge evaluates the generalization capabilities of automated mitosis detection models across diverse tumor types, species, and scanning platforms, revealing significant performance degradation in challenging real-world contexts and highlighting the limitations of current architectures despite the benefits of ensembling.

Original authors: Marc Aubreville, Jonas Ammeling, Sweta Banerjee, Viktoria Weiss, Taryn A. Donovan, Robert Klopfleisch, Jiaqi Lv, Shan E Ahmed Raza, Raphaël Bourgade, Thomas Walter, Yasemin Topuz, Songül Varlı, Charle
Published 2026-06-08
📖 5 min read🧠 Deep dive

Original authors: Marc Aubreville, Jonas Ammeling, Sweta Banerjee, Viktoria Weiss, Taryn A. Donovan, Robert Klopfleisch, Jiaqi Lv, Shan E Ahmed Raza, Raphaël Bourgade, Thomas Walter, Yasemin Topuz, Songül Varlı, Charles-Antoine Collins-Fekete, Zhuoyan Shen, Navya Sri Kelam, Nitin Singhal, Christian Marzahl, Brian Napora, Tengyou Xu, Hongyan Gu, Mario Vento, Gennaro Percannella, Norbert Ropiak, Izabela Wasiak, Jie Xiao, Shaojun Liu, Seungho Choe, April Khademi, Vidushi Walia, Sujatha Kotte, Andrew Broad, Alex Wright, Guillaume Balezo, Esha Sadia Nasir, Mostafa Jahanifar, Yosuke Yamagishi, Shouhei Hanaoka, Mattia Sarno, Francesco Tortorella, Biwen Meng, Jingxin Liu, Sara Krauss, Daniel Hieber, Lavish Ramchandani, Dev Kumar Das, Mieko Ochi, Yuan Bae, Piotr Giedziun, Mateusz Maniewski, Vangala Govindakrishnan Saipradeep, Naveen Sivadasan, Leire Benito-Del-Valle, Adrian Galdran, Kaustubh Atey, Sameer Anand Jha, Adinath Dukre, Imran Razzak, Maxime W. Lafarge, Viktor H. Koelzer, Nils Porsche, Nikolas Stathonikos, Mitko Veta, Dominik Hirling, Zsanett Zsófia Iván, Peter Horvath, Katharina Breininger, Christof A. Bertram

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to spot a specific type of tiny, busy worker (a dividing cell) inside a massive, chaotic city (a tissue sample). For years, researchers only taught these robots by showing them photos of the city's busiest, most obvious construction zones. The robots got really good at finding workers in those specific zones.

But in the real world, doctors don't just look at the busiest zones; they have to scan the entire city, including quiet neighborhoods and confusing construction sites that look like workers but aren't.

This paper describes the MIDOG 2025 Challenge, a massive competition where 18 teams of computer scientists tried to build robots that could find these "workers" (mitotic figures) anywhere in the city, not just in the easy spots. They also had a second task: figuring out if a worker they found was doing their job normally or if they were "atypical" (doing something weird and potentially dangerous).

Here is a breakdown of what they found, using simple analogies:

1. The "Spot the Worker" Game (Track 1)

The first task was to find every dividing cell in three different types of neighborhoods:

  • The "Hotspot" Neighborhood: The obvious, busy construction zone where workers are everywhere. (This is what previous challenges used).
  • The "Random" Neighborhood: A random slice of the city, representing a normal, average view of the whole tissue.
  • The "Challenging" Neighborhood: A tricky area filled with "imposters"—things that look like workers (like dead cells or ink marks) but aren't.

The Big Surprise:
The robots were excellent at the "Hotspot" neighborhoods. But as soon as they stepped into the "Random" or "Challenging" neighborhoods, they started making mistakes.

  • The Analogy: Imagine a security guard who is perfect at spotting thieves in a crowded stadium (Hotspot). But when you put that same guard in a quiet library (Random) or a room full of people wearing identical costumes (Challenging), they start screaming "Thief!" at innocent people.
  • The Result: In the "Challenging" areas, the robots' false alarms (calling innocent things "workers") jumped by 208%. This means the current technology isn't ready to scan a whole tissue slide reliably yet; it's too prone to panic in confusing areas.

2. The "Normal vs. Weird" Game (Track 2)

The second task was to look at the workers the robots did find and decide: "Is this a normal worker, or is this a weird, dangerous one?"

  • The Result: The robots were much better at this. The top teams got about 91% accuracy.
  • The Catch: Even here, the robots struggled with rare types of tumors. It's like having a security guard who is great at spotting weird behavior in a crowd of 10,000 people, but gets confused when the crowd is made of a very specific, rare group of people they've never seen before.

3. The "Secret Weapons" (What actually helped?)

The researchers peeked under the hood of the winning robots to see what tricks they used.

  • The "Group Think" Strategy (Ensembling):

    • What it is: Instead of using one robot, some teams used a team of 3 to 5 robots and let them vote on the answer.
    • Did it work? Yes. This was the most consistent booster. It was like having a committee of experts instead of a single opinion. It improved scores across the board, especially in the tricky areas.
    • Note: The absolute #1 winner in the second track didn't use this trick, proving you don't need a committee to win, but it usually helps.
  • The "Double-Check" Strategy (Test-Time Augmentation):

    • What it is: This involves showing the robot the same image multiple times, but flipped, rotated, or zoomed, to see if it still spots the worker.
    • Did it work? No. It didn't really help. In fact, it just made the robot slower and used more computer power without making it smarter. It's like asking a person to look at a photo, then turn the photo upside down and look again, hoping they'll find a mistake they missed the first time. For these robots, it was just a waste of time.

4. The "Blind Spots"

The paper highlights that the robots have "blind spots."

  • Tumor Types: The robots worked well on common tumors but struggled significantly with rare or very messy-looking tumors (like human glioblastoma).
  • Context: The robots were trained mostly on small, cropped images of just the worker. They missed the "big picture" of the surrounding neighborhood. The paper suggests that to get better, future robots need to see the whole neighborhood, not just the worker's face.

The Bottom Line

The MIDOG 2025 challenge proved that while AI is getting very good at finding dividing cells in the "easy" parts of a tissue slide, it is still not reliable enough for the whole slide yet.

  • Current State: Great at the "Hotspots," terrible at the "Challenging" areas.
  • The Fix: We need to train these robots on more diverse, messy, and "real-world" data, not just the perfect, curated photos.
  • The Takeaway: Before we trust these robots to scan a patient's entire tissue sample in a hospital, we need to make sure they don't get confused by the "imposters" in the challenging neighborhoods.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →