Comparative Performance of General-Purpose Large Language Models and a Specialized Clinical AI Tool on OKAP-Style Ophthalmology Questions
This study evaluates five large language models on OKAP-style ophthalmology questions, revealing that general-purpose models, particularly Gemini-Pro-3, outperformed the specialized clinical AI tool OpenEvidence in accuracy and calibration, while highlighting significant performance heterogeneity and shared reasoning failures across all systems.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a high-stakes medical exam for eye doctors, known as the OKAP. Now, imagine five different "super-smart" computer programs trying to take this test. The goal of this study was to see which computer program is the best at answering these tricky eye-doctor questions and, just as importantly, which one knows when it might be wrong.
Here is the breakdown of what happened, using some everyday analogies:
The Contestants
The researchers lined up five different AI "students":
- OpenEvidence: A specialized student who only studies medical textbooks and has a direct line to famous medical journals. Think of this as a medical librarian who has memorized the library but doesn't have a general brain for everything else.
- Gemini-Pro-3, Claude-Opus-4.6, ChatGPT-5.4-Pro, and DeepSeek-v3.2: These are "general-purpose" students. They read everything from cooking recipes to physics papers. Think of them as polymaths—people who know a little bit about a lot of things.
The Test
The test consisted of 659 multiple-choice questions about eye health. These weren't easy trivia; they ranged from simple facts (like "What is the name of this eye part?") to complex scenarios requiring deep reasoning (like "If a patient has symptom X and history Y, what is the best treatment?").
The Results: Who Won?
The results were surprising to some experts:
- The Champion: Gemini-Pro-3 (the general-purpose student) got the highest score, answering 95.8% of the questions correctly. It was so good that it beat everyone else by a significant margin.
- The Runner-Up: Claude-Opus-4.6 (another general-purpose student) came in second.
- The Specialist: OpenEvidence (the medical librarian) came in third. It did very well (88%), but it did not beat the general-purpose students. In fact, the general students outperformed the specialist.
- The Rest: ChatGPT and DeepSeek followed, with DeepSeek scoring the lowest (73.7%).
The Big Takeaway: Having a tool built specifically for medicine didn't guarantee it would be the smartest at answering medical questions. The general-purpose AI was actually better at this specific test.
The "Confidence" Check
The researchers also asked the AI: "How sure are you that your answer is right?" (On a scale of 0 to 100). This is like asking a student to raise their hand if they are 100% confident.
- The Best "Self-Aware" Student: Gemini-Pro-3 was the most honest. When it said it was confident, it was usually right. Its confidence matched its actual performance perfectly.
- The "Over-Confident" vs. "Under-Confident": Some other models were good at knowing when they were right (discrimination), but not as good at matching their confidence level to their actual accuracy (calibration).
- The Warning Sign: One model, DeepSeek, was terrible at this. It often said it was confident when it was actually wrong, or unsure when it was right. This is dangerous because if a doctor trusts a confident but wrong answer, it could lead to mistakes.
Where Did They All Fail?
The researchers looked at the 9 questions that every single AI got wrong. They tried to figure out why.
- What they got right: They understood the words and the basic topic.
- Where they broke down: They failed at complex reasoning and checking their sources.
- Analogy: Imagine a student who can read a question perfectly but gets lost when they have to connect three different facts to solve a puzzle. Or, they might guess an answer without actually checking if the rulebook supports it.
- The AI struggled most with questions that required "thinking steps" rather than just remembering a fact.
The Limitations (The Fine Print)
The authors were careful to say what this study didn't prove:
- No Pictures: The test was only text. In real eye care, doctors look at photos of eyes. None of these AIs were tested on looking at images during this study.
- Not Real Patients: These were exam questions, not real-life patients with messy, complicated histories.
- One-Time Test: The AI took the test once. If you asked it again, it might give a different answer.
Summary
In simple terms: General-purpose AI is currently beating specialized medical AI at answering eye-doctor exam questions. However, the AI still struggles with the hardest, most complex reasoning tasks. Also, just because an AI gives a high score doesn't mean it's safe to use blindly; you need to check if it knows when it's unsure. The study suggests that for now, we should be careful about using these tools for real medical decisions until we know more about how they handle real-world complexity.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.