← Latest papers
🤖 AI

Deep Learning for Retinal Degeneration Assessment: A Comprehensive Analysis of the MARIO Challenge

This paper presents a comprehensive analysis of the MARIO challenge at MICCAI 2024, detailing how 35 teams utilized multi-modal OCT and clinical data to benchmark AI performance in retinal degeneration assessment, revealing that while AI matches physician accuracy in detecting AMD progression, it currently falls short in predicting future disease evolution.

Original authors: Rachid Zeghlache, Ikram Brahim, Pierre-Henri Conze, Mathieu Lamard, Mohammed El Amine Lazouni, Zineb Aziza Elaouaber, Leila Ryma Lazouni, Christopher Nielsen, Ahmad O. Ahsan, Matthias Wilms, Nils D. F
Published 2026-08-04
📖 6 min read🧠 Deep dive

Original authors: Rachid Zeghlache, Ikram Brahim, Pierre-Henri Conze, Mathieu Lamard, Mohammed El Amine Lazouni, Zineb Aziza Elaouaber, Leila Ryma Lazouni, Christopher Nielsen, Ahmad O. Ahsan, Matthias Wilms, Nils D. Forkert, Lovre Antonio Budimir, Ivana Matovinović, Donik Vršnak, Sven Lončarić, Philippe Zhang, Weili Jiang, Yihao Li, Yiding Hao, Markus Frohmann, Patrick Binder, Marcel Huber, Taha Emre, Teresa Finisterra Araújo, Marzieh Oghbaie, Hrvoje Bogunović, Amerens A. Bekkers, Nina M. van Liebergen, Hugo J. Kuijf, Abdul Qayyum, Moona Mazher, Steven A. Niederer, Alberto J. Beltrán-Carrero, Juan J. Gómez-Valverde, Javier Torresano-Rodríquez, Álvaro Caballero-Sastre, María J. Ledesma Carbayo, Yosuke Yamagishi, Yi Ding, Robin Peretzke, Alexandra Ertl, Maximilian Fischer, Jessica Kächele, Sofiane Zehar, Karim Boukli Hacene, Thomas Monfort, Béatrice Cochener, Mostafa El Habib Daho, Anas-Alexis Benyoussef, Gwenolé Quellec

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery that happens inside a tiny, intricate city: the human eye. Specifically, you are looking at the retina, the light-sensitive screen at the back of the eye. Sometimes, this city gets invaded by a sneaky, leaky growth called neovascular age-related macular degeneration (AMD). It's like a plumbing disaster where tiny, unwanted pipes sprout and leak fluid, blurring the center of your vision. To fix this, doctors use special "anti-leak" injections (anti-VEGF) that act like a super-glue to stop the pipes. But here's the tricky part: the leakiness changes every day. Sometimes the pipes are quiet, sometimes they are gushing, and sometimes they are just a little damp. Doctors have to take high-resolution 3D pictures of the eye (called OCT scans) every few weeks to see if the glue is working. If they guess wrong, they might inject a patient who doesn't need it, or worse, wait too long on a patient who is getting worse.

For a long time, computers have been getting really good at looking at a single picture and saying, "Hey, this eye has a leak!" But the real challenge is like watching a movie instead of a still photo. The computer needs to look at two pictures taken weeks apart and decide: "Did the leak get better, stay the same, or get worse?" Even harder, can the computer look at just one picture today and predict what will happen three months from now? This is the question a group of scientists asked the world's best AI detectives to solve in a global contest called the MARIO challenge. They wanted to know if artificial intelligence could become a reliable sidekick for doctors, helping them decide exactly when to give those injections and when to let the patient rest.

The Great Eye Detective Contest

In 2024, a team of researchers organized a massive game called the MARIO challenge (Monitoring Age-Related macular degeneration with Intelligent Ophthalmology). They gathered 35 teams of AI experts from around the world, including students, university researchers, and tech companies. The goal was to build a computer program that could analyze optical coherence tomography (OCT) scans—those super-detailed 3D cross-sections of the retina—to track the disease.

The challenge had two levels, like a video game with an easy mode and a hard mode.

Level 1: The "Spot the Change" Game
In this level, the AI was given two pictures of the same eye taken at different times (let's say, Visit A and Visit B). The computer had to look at both and decide: Did the disease get Reduced (better), stay Stable (the same), or Worsened (got worse)?
The results here were surprisingly exciting. The best team, called MIPLAB, built a model that got it right about 86% of the time. To put that in perspective, this is almost as good as two human eye doctors agreeing with each other! The AI successfully learned to spot the subtle differences in fluid levels between two scans. It was like teaching a computer to notice that a puddle on the sidewalk had shrunk or grown just by looking at two photos. This suggests that for checking if a treatment is working right now, AI is ready to be a helpful assistant.

Level 2: The "Crystal Ball" Game
This level was much, much harder. The AI was given only one picture (Visit A) and had to predict what would happen in the next three months. Would the eye get better, stay the same, or get worse?
The results here were a reality check. The best AI model only managed a score of about 0.30, which is barely better than just guessing "Stable" every time because most eyes are stable. The teams' predictions were so different from each other that they agreed with one another no better than chance. It was as if the AI was trying to predict the weather three months in advance using only a single photo of the sky today; there just wasn't enough information in that one picture to know what was coming. The paper concludes that predicting the future from a single visit is still a problem that AI hasn't solved yet.

The "Foreign Language" Problem

There was one more twist in the story that taught the researchers a very important lesson about fairness and reliability. The main group of patients came from Brest, France. But to test if the AI was truly smart or just memorized the French data, the organizers threw in a tiny, secret test set from Tlemcen, Algeria. These patients had different backgrounds, and their eye scans were taken with slightly different equipment settings.

When the winning AI from the French data tried to solve the Algerian puzzles, it stumbled badly. The team that came in first place in France dropped to fifth place in Algeria. Some teams' scores crashed so hard they were barely better than random guessing. It was like a student who aced a math test in one classroom but failed completely when the teacher switched to a different textbook with slightly different symbols. This showed that the AI models were too sensitive to where the data came from. They hadn't learned the universal rules of the eye; they had just learned the specific "accent" of the French scans.

What Does This Mean?

The MARIO challenge gave us a clear map of where we stand.

  1. Short-term monitoring is possible: If a doctor wants to know if an eye is better or worse compared to last month's scan, AI can do a great job. It's close to being ready to help doctors save time and reduce unnecessary injections.
  2. Long-term prediction is not ready: If a doctor wants to know what will happen three months from now based on a single scan, AI is currently useless. The signal is too weak, and the models are just guessing.
  3. One size does not fit all: An AI that works perfectly for one group of people might fail miserably for another. Before we can trust these tools in hospitals everywhere, we need to make sure they work for everyone, regardless of where they live or what machine scanned their eyes.

The paper doesn't claim to have solved the mystery of AMD, but it did solve the mystery of what AI can and cannot do right now. It tells us that while we have built a very good "spotter" for current changes, we still need to invent a much smarter "prophet" for the future. And until we do, we must be careful not to trust our digital detectives too much when they are in a new neighborhood.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →