← Latest papers
💻 computer science

RetiBridge: Bridging Quantitative Retinal Biomarkers and Qualitative Diagnosis with a Knowledge-Guided Multimodal Large Language Model

RetiBridge is a knowledge-guided multimodal large language model that effectively bridges quantitative retinal biomarkers from CFP and OCT images with qualitative clinical diagnoses, outperforming existing open-source and proprietary models on a new benchmark of over 15,000 paired samples.

Original authors: Zhuangzhi Gao, Hongyi Qin, He Zhao, Qinkai Yu, Feixiang Zhou, Fu Wang, Jinru Ding, Eduard Shantsila, Uazman Alam, Alena Shantsila, Wahbi El-Bouri, Gregory Y. H. Lip, Yalin Zheng

Published 2026-07-31
📖 4 min read☕ Coffee break read

Original authors: Zhuangzhi Gao, Hongyi Qin, He Zhao, Qinkai Yu, Feixiang Zhou, Fu Wang, Jinru Ding, Eduard Shantsila, Uazman Alam, Alena Shantsila, Wahbi El-Bouri, Gregory Y. H. Lip, Yalin Zheng

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine your eyes are like a high-tech security camera system for your entire body. Just as a security camera can spot a broken window or a strange shadow, the retina at the back of your eye can reveal clues about diseases like diabetes, high blood pressure, and even heart issues. Doctors use two main "lenses" to look at this security feed: one takes a flat, colorful photo of the surface (like a standard snapshot), and the other takes a deep, 3D slice through the layers (like a cross-section of a cake). For a long time, computers were great at measuring the numbers in these pictures—like how thick a layer is or how wide a blood vessel is—but they were terrible at explaining what those numbers actually mean in plain English. It was like having a robot that could tell you a car's speed is 60 mph but couldn't tell you if that meant the car was speeding or just driving on a highway.

This is where the new research comes in. Scientists have been trying to build "Multimodal Large Language Models" (MLLMs)—basically super-smart AI brains that can see images and read text at the same time. The goal is to create a digital doctor's assistant that doesn't just say "abnormal," but explains why by connecting the hard numbers to a real diagnosis. However, most existing AI models are like students who memorized the textbook but can't apply it to a real-life exam; they might guess the right disease but can't show their work or explain how the measurements led to that conclusion. This paper introduces a new AI called RetiBridge that tries to fix this by forcing the computer to act like a careful detective, linking every single number it finds to a specific medical clue before making a final verdict.

The researchers built RetiBridge to be a "knowledge-guided" detective. Instead of letting the AI guess, they taught it a specific way of thinking: first, measure the evidence (the numbers); second, translate those numbers into a medical story (the clues); and third, solve the case (the diagnosis). They trained this AI using 15,611 pairs of eye images from the UK Biobank, a massive database of real patient data. For every patient, the AI had access to 31 different measurements from the deep 3D scans and 6 measurements from the flat photos. To make sure the AI learned the right way to think, the researchers used a powerful AI tool (OpenAI-o3) to write "model answers" for these 15,000+ cases. These model answers didn't just give a diagnosis; they wrote out the full reasoning, showing exactly how a measurement like "240 micrometers of thickness" translates to "no swelling here," which then leads to a final conclusion.

When they tested RetiBridge, it showed some impressive skills. Even though it was built on a relatively small "brain" (a 7-billion-parameter model), it outperformed much larger models, including some that are 32 billion parameters and even a very advanced proprietary model from OpenAI. The key to its success was a special training step where the AI learned to align the 3D eye images directly with the specific numbers they represent, kind of like teaching a student to match a picture of a cake slice directly to a ruler, rather than just memorizing the word "cake."

The results suggest that RetiBridge is better at the specific job of "showing its work." It was significantly more accurate at getting the numbers right (78.23% accuracy compared to much lower scores for other models) and better at proving that its diagnosis was actually supported by the evidence it saw. While other models sometimes got the final diagnosis right but missed the details or made up reasons, RetiBridge consistently linked its conclusions back to the actual measurements. For example, in a test case involving diabetic retinopathy, RetiBridge correctly identified specific spots and thicknesses and explained how they ruled out other diseases, creating a clear, logical path from the photo to the answer.

However, the paper is careful to note that this isn't a magic cure-all yet. While RetiBridge is better at connecting the dots, it still struggles with some complex reasoning tasks compared to the very best human-like models, and it hasn't been tested on every single type of eye disease. The authors suggest that this approach—forcing AI to ground its answers in hard numbers—is a promising path forward, but more work is needed to make it robust enough for real-world hospitals. They also plan to test it on even more detailed 3D scans and different types of eye images in the future. For now, RetiBridge stands as a strong proof-of-concept that if you teach an AI to respect the math behind the medicine, it can become a much more trustworthy partner in understanding our health.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →