Can Multimodal Large Language Models Understand OCT?
This paper introduces OCT-Bench, a comprehensive benchmark comprising over 10,000 fine-grained questions across perception, cognition, and reasoning dimensions to systematically evaluate multimodal large language models, revealing that current models—including medical-adapted and larger-scale variants—remain significantly limited in their ability to perform clinically grounded OCT image understanding.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery, but instead of a crime scene, your evidence is a glowing, cross-sectional map of a human eye. This is the world of Optical Coherence Tomography (OCT), a super-powered camera that lets doctors peek inside the retina—the light-sensitive layer at the back of the eye—without making a single cut. It's like having an X-ray vision for the eye's tiny, layered neighborhoods, helping doctors spot diseases before they steal your sight. Now, enter the new heroes of the tech world: Multimodal Large Language Models (MLLMs). Think of these as super-smart AI robots that can read text and look at pictures at the same time. They are the "Sherlock Holmes" of the digital age, trained to connect what they see with what they know. But here is the big question that has everyone buzzing: Can these AI detectives actually understand the complex, layered maps of the human eye, or are they just guessing based on patterns? If an AI can't tell the difference between a healthy eye and a sick one, or can't figure out the right treatment, it's not much use to a doctor. This is why scientists are so eager to test if these models are truly ready to help save vision.
In this paper, a team of researchers decided to stop guessing and start testing. They built a massive, super-challenging exam called OCT-Bench to see if AI can really handle the job of reading eye scans. Instead of just asking the AI, "Is this eye sick or healthy?" (which is like asking a student to just guess the answer on a multiple-choice test), they broke the task down into a step-by-step journey that mirrors how a real doctor thinks. First, the AI has to Perceive: Can it see the basic shapes, colors, and lines in the image? Next, it must Cognize: Can it recognize that a specific squiggly line is actually a layer of the retina or a fluid-filled pocket? Finally, it has to Reason: Can it take all those clues and decide what disease it is, what treatment to give, and what will happen next?
To make this exam fair and tough, the researchers gathered 4,137 real eye scans from seven different public databases and turned them into 10,076 expert-verified questions. They then put 20 different AI models through the wringer, including the famous "big brains" from tech giants, open-source models, and special AI trained just for medicine.
The results? The AI models are definitely smart, but they are not yet the eye-doctors of the future. The best-performing model in the entire test only got 62.0% of the answers right. That means it was wrong on nearly 4 out of every 10 questions. The researchers found a clear pattern: the AI was okay at the easy stuff. When asked to simply identify the type of image or count how many boxes were drawn on the screen, the models scored high, sometimes near 75.8%. But as soon as the questions got harder—requiring the AI to understand the tiny details of the eye's layers or to figure out a treatment plan—the scores crashed. For the hardest "Reasoning" tasks, the best model only scored 42.9%.
Here is the surprising part: being a "medical expert" or having a bigger brain didn't automatically fix the problem. Some models trained specifically on medical data did worse than general-purpose models on certain tasks. And making the models bigger (scaling them up) helped a little bit, but it didn't solve the big issue. The AI still struggles to connect what it sees in the picture with the deep medical knowledge needed to make a real diagnosis. It's like a student who can memorize the names of all the bones in the body but freezes when asked to diagnose a broken leg based on an X-ray.
The paper concludes that current AI models remain substantially short of reliable OCT understanding and are far from reliable for clinical use. It suggests that while these models are getting better at seeing, they are still missing the "clinical reasoning" muscle needed to understand why something is wrong and what to do about it. The researchers indicate that until AI can reliably move from simply "seeing" the image to "understanding" the disease and "planning" the cure, we have a long way to go before we can trust them with our eyes.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.