← Latest papers
🤖 AI

Full end-to-end diagnostic workflow automation of 3D OCT via foundation model-driven AI for retinal diseases

The paper introduces FOCUS, a foundation model-driven framework that automates the end-to-end diagnostic workflow for 3D OCT retinal imaging by integrating quality assessment, abnormality detection, and multi-disease classification into a unified system that achieves expert-level accuracy and efficiency across diverse clinical settings.

Original authors: Jinze Zhang, Jian Zhong, Li Lin, Jiaxiong Li, Ke Ma, Naiyang Li, Meng Li, Yuan Pan, Zeyu Meng, Mengyun Zhou, Shang Huang, Shilong Yu, Zhengyu Duan, Sutong Li, Honghui Xia, Juping Liu, Dan Liang, Yanta
Published 2026-08-25
📖 6 min read🧠 Deep dive

Original authors: Jinze Zhang, Jian Zhong, Li Lin, Jiaxiong Li, Ke Ma, Naiyang Li, Meng Li, Yuan Pan, Zeyu Meng, Mengyun Zhou, Shang Huang, Shilong Yu, Zhengyu Duan, Sutong Li, Honghui Xia, Juping Liu, Dan Liang, Yantao Wei, Xiaoying Tang, Jin Yuan, Peng Xiao

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The human eye is a window to the body, but looking through it requires a steady hand and a trained eye. For decades, doctors have relied on a scanning technology called optical coherence tomography, or OCT, to peer inside the retina, the light-sensitive layer at the back of the eye. This machine takes thousands of cross-sectional pictures, stacking them to create a detailed three-dimensional map of the eye's internal structure. It has revolutionized how doctors spot diseases like diabetic retinopathy or macular degeneration, allowing them to see problems long before a patient loses vision. Yet, despite the machine's ability to capture this rich data, the process of turning those images into a diagnosis remains a slow, manual chore. A specialist must first check if the scan is clear enough to use, then sift through hundreds of individual slices to find the few that show signs of trouble, and finally piece those fragments together to form a complete picture of the patient's health. This labor-intensive workflow creates a bottleneck, making it difficult to screen large populations or provide consistent care in places where expert doctors are scarce.

Researchers have long tried to build artificial intelligence to help with this task, but most previous attempts have been like a team of specialists who only know how to do one small job. One computer program might be good at spotting blurry images, while another is excellent at identifying a specific disease on a single slice of the scan. None of these systems could handle the entire journey from raw image to final diagnosis on their own. They often struggled with the fact that eye diseases are three-dimensional, hiding in the space between slices, and they faltered when faced with the messy, varied reality of real-world hospital data. A new study introduces a system designed to change this by automating the entire process, mimicking the way a human expert thinks through a case from start to finish.

The team behind this work, led by researchers from several major medical centers in China, developed a system they call FOCUS. This framework is built on a foundation model, a type of artificial intelligence that has been pre-trained on a vast amount of general visual data, giving it a broad understanding of shapes and patterns before it ever sees an eye scan. The researchers fine-tuned this powerful brain to understand retinal images, but the true innovation lies in how the system is organized. Instead of just looking at one picture at a time, FOCUS acts as a complete workflow. First, it automatically checks the quality of the scan, deciding if the image is clear enough to trust. If the image is good, it moves to the next step: scanning through the entire volume of the eye to detect any abnormalities. Finally, it synthesizes all the information from the thousands of slices to make a single, comprehensive diagnosis for the patient.

To test if this approach worked, the researchers trained the system on data from 3,300 patients, a massive collection of over 40,000 individual image slices. They then put the system to the test on a separate group of 1,345 patients from four different hospitals, using different types of OCT machines to ensure it could handle real-world variety. The results were striking. The system correctly identified high-quality images nearly 99% of the time. When looking for signs of disease, it caught abnormalities with an accuracy of over 97%. Most importantly, when it came to making a final diagnosis for the patient, it achieved an accuracy of nearly 94%. These numbers held up even when the system was tested on data from different cities and different hospitals, showing that it could adapt to new environments without needing to be retrained from scratch.

The researchers also compared the performance of their AI against human doctors to see how they stacked up. In tests involving the detection of abnormalities, the system performed as well as, and in some cases better than, a group of experienced ophthalmic technicians. When it came to diagnosing specific diseases, the system matched the performance of a panel of nine clinicians with varying levels of experience. In fact, the AI was more consistent than the humans. The study noted that even expert doctors sometimes struggled to distinguish between certain conditions when looking at OCT images alone, often needing to rely on other tests or patient history to be sure. The AI, however, was able to spot subtle structural differences across the entire 3D volume of the eye that humans might miss, leading to more uniform decisions.

A key part of why this system succeeds is how it handles the three-dimensional nature of the eye. Previous AI tools often treated each slice of the scan as an isolated event, missing the context that connects them. FOCUS uses a method that allows it to weigh the importance of each slice dynamically. It learns to focus on the specific slices that contain the strongest evidence of disease while ignoring the ones that are ambiguous or irrelevant. This is similar to how a human expert might scan a stack of papers, quickly glancing through the boring ones and stopping to read carefully only the pages that contain the critical information. By aggregating these findings intelligently, the system builds a reliable diagnosis without needing to process the entire 3D volume in a computationally expensive way.

The study acknowledges that while the results are promising, the system is not yet a finished product ready for every clinic. The data used to train and test the system came primarily from medical centers in China, which means the system might need further validation to ensure it works equally well for people of different ethnic backgrounds or in different healthcare systems. Additionally, the testing was done on data collected from hospitals where patients were already known to have eye problems, which is a different scenario than screening a healthy general population where diseases are rarer. The researchers also point out that the current version of the system does not yet include the ability to generate written reports or hold a conversation with a doctor, features that future versions might add using advanced language tools.

Despite these limitations, the work represents a significant step forward in the field of medical imaging. It moves beyond the era of isolated tools that perform single tasks and demonstrates that a unified system can automate a complex, multi-step clinical workflow. By successfully bridging the gap between raw image data and a final patient diagnosis, the researchers have shown that it is possible to build an AI that mirrors the reasoning of a human specialist. This approach offers a blueprint for a future where high-quality eye care can be delivered more efficiently and consistently, potentially bringing expert-level screening to communities that currently lack access to specialized doctors. The system does not replace the doctor but rather removes the tedious, repetitive barriers that prevent large-scale screening from becoming a reality, paving the way for a new era of automated, accessible eye health.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →