← Latest papers
💻 computer science

Development and evaluation of CADe systems in low-prevalence setting: The RARE25 challenge for early detection of Barrett's neoplasia

The RARE25 challenge establishes a prevalence-aware benchmark for early Barrett's neoplasia detection, revealing that while current CADe systems achieve strong discriminative performance, their low positive predictive values in realistic low-prevalence settings highlight the critical need for robust, prevalence-agnostic approaches to ensure clinical utility.

Original authors: Tim J. M. Jaspers, Francisco Caetano, Cris H. B. Claessens, Carolus H. J. Kusters, Rixta A. H. van Eijck van Heslinga, Floor Slooter, Jacques J. Bergman, Peter H. N. De With, Martijn R. Jong, Albert J
Published 2026-04-15
📖 6 min read🧠 Deep dive

Original authors: Tim J. M. Jaspers, Francisco Caetano, Cris H. B. Claessens, Carolus H. J. Kusters, Rixta A. H. van Eijck van Heslinga, Floor Slooter, Jacques J. Bergman, Peter H. N. De With, Martijn R. Jong, Albert J. de Groof, Fons van der Sommen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a security guard at a massive airport. Your job is to scan thousands of passengers every day to find the one person who might be carrying a dangerous weapon.

Most of the time, everyone is innocent. You see a sea of friendly faces, suitcases, and normal people. But if you miss that one person, the consequences are terrible. If you stop too many innocent people, you cause chaos, anger, and "alarm fatigue" (where you stop caring because you're stopped so many false alarms).

This is exactly the problem doctors face when looking for early cancer in the esophagus (a condition called Barrett's Esophagus). They have to scan thousands of images of healthy tissue to find a tiny, subtle spot of cancer.

This paper describes a giant experiment called RARE25 (Recognition of Abnormalities in low-pREvalence cancer) designed to solve this problem using Artificial Intelligence (AI).

Here is the story of the challenge, the contestants, and what they learned, explained simply.

1. The Problem: The "Needle in a Haystack" Effect

For years, AI researchers have been building "Computer-Aided Detection" (CADe) systems to help doctors find these tiny cancer spots. But there was a big lie in how they tested these AI systems.

  • The Fake Test: Researchers usually trained their AI on datasets where they artificially balanced the numbers. They might show the AI 50 pictures of cancer and 50 pictures of healthy tissue. It's like training a security guard by showing them 50 bad guys and 50 good guys. The AI gets really good at spotting the bad guys in this fake world.
  • The Real World: In reality, out of 10,000 patients, maybe only 100 have early cancer. That is a 1% prevalence.
  • The Result: When these "perfect" AIs were deployed in the real world, they went crazy. Because they were trained to expect cancer often, they started screaming "CANCER!" at every tiny speck of dust. They created so many false alarms that doctors couldn't trust them.

2. The Challenge: RARE25

To fix this, the organizers created the RARE25 Challenge. They wanted to see if AI could handle the real difficulty of the job.

  • The Training Set: They gave the teams a large library of images (about 3,000), but it was still mostly healthy tissue with very few cancer spots.
  • The Secret Test: They held back a massive "secret test" of over 25,000 images. Crucially, this test set was sampled to match real life. It had the same tiny ratio of cancer to healthy tissue that doctors actually see in hospitals.
  • The Goal: The teams had to build an AI that could find the cancer without screaming "False Alarm" at everything else.

3. The Contestants: 11 Teams from Around the World

Eleven teams from countries like Germany, Japan, South Korea, and the Netherlands entered the arena. They used different strategies, like different types of "brains" (neural networks) and different training tricks.

Here are a few of their approaches:

  • The "Super-Ensemble" Team (IMSY): Instead of using one smart AI, they built a team of 40 different AIs and asked them all to vote. They also used a "foundation model" (a pre-trained brain that already knows a lot about medical images) and fine-tuned it. They won the competition!
  • The "Segmentation" Team (Jmees-inc): They tried a different angle. Instead of just saying "Yes/No," they tried to draw a box around the suspicious area first, then decide if it was cancer.
  • The "Pre-training" Teams: Many teams tried to teach their AI on other medical datasets (like polyps in the colon) before tackling this specific cancer, hoping the AI would learn general medical skills first.

4. The Results: Good News and Bad News

When the results came in, it was a mix of victory and reality checks.

  • The Good News: The AI systems were incredibly good at ranking images. If you showed them a cancer and a healthy tissue, they could almost always tell which was which. They had high "discrimination."
  • The Bad News (The Reality Check): When they had to actually detect the cancer in the real-world test set (where cancer is rare), the Positive Predictive Value (PPV) was still very low.
    • Translation: Even the winning team, when they said "This is cancer," they were only right about 3.5% of the time. The other 96.5% were false alarms.
    • Why? Because the cancer is so rare, even a tiny mistake rate creates a mountain of false alarms.

The Analogy: Imagine the AI is a metal detector. It is so sensitive that it beeps for a coin, a belt buckle, and a piece of foil. It can find the gun, but it also beeps at everything else. In a crowd of 10,000 people, if it beeps 1,000 times, the security guard has to check 1,000 innocent people to find the one bad guy. That's not practical yet.

5. What Did We Learn? (The "Aha!" Moments)

The paper highlights three major lessons for the future:

  1. Stop Faking the Test: You cannot train and test AI on balanced datasets. You must test it on data that looks like the real world, or the results are meaningless.
  2. The "False Alarm" Trap: In low-prevalence settings, being "accurate" isn't enough. If an AI creates too many false alarms, it becomes useless because it overwhelms the doctor.
  3. The Missing Strategy: Almost every team tried to teach the AI to recognize "Cancer." But nobody tried to teach the AI to recognize "Normal."
    • The Future Idea: Instead of teaching the AI what cancer looks like (which is rare), maybe we should teach it what healthy looks like (which is everywhere). Then, if the AI sees something that doesn't look healthy, it raises an alarm. This is called "Anomaly Detection," and it might be the key to solving this problem.

Conclusion

The RARE25 challenge was a success, not because the AI solved the problem perfectly, but because it exposed the truth. It showed us that while AI is getting smarter at spotting patterns, we still have a long way to go before it can reliably act as a "second pair of eyes" in a real hospital without causing a panic of false alarms.

It's a reminder that in medicine, context is everything. An AI that works in a lab might fail in a hospital, and the only way to know is to test it under real-world pressure.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →