Evidence-Grounded Subspecialty Reasoning: Evaluating a Curated Clinical Intelligence Layer on the 2025 Endocrinology Board-Style Examination
The study demonstrates that January Mirror, an evidence-grounded clinical reasoning system with a curated endocrinology corpus, significantly outperformed both frontier large language models with real-time web access and human experts on a 2025 endocrinology board-style examination, achieving 87.5% accuracy while ensuring 100% citation verifiability.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to solve a very tricky medical mystery. You have two detectives on the case:
- The "Web Surfer" Detective: This detective is incredibly smart and has access to the entire internet. They can read millions of articles, blogs, and news stories in seconds. However, they have to figure out which stories are true, which are outdated, and which are just noise. They might get distracted by a flashy headline or accidentally mix up an old rule with a new one.
- The "Librarian" Detective (Mirror): This detective doesn't have the internet. Instead, they have a single, perfectly organized, ultra-secure library. Inside this library, every book has been hand-picked by the world's top experts. The books are arranged by importance: the most important rules are on the top shelf, and the supporting evidence is neatly filed below. This detective is trained to only use these specific books to solve the case.
The Big Test
The authors of this paper put both detectives to the test using the ESAP 2025, which is like the "Olympics" for endocrinologists (doctors who specialize in hormones, diabetes, and metabolism). It's a 120-question exam that is notoriously difficult, filled with complex patient stories that require deep, specialized thinking.
The Results: The Librarian Wins
Here is the surprising twist: The Librarian (Mirror) won by a landslide.
- The Librarian (Mirror): Got 87.5% of the questions right.
- The Web Surfer (Top AI models like GPT-5.2): Got about 74.6% right.
- The Average Human Doctor: The average human taking this test got about 62.3% right.
Even though the "Web Surfer" had the whole internet at its fingertips, the "Librarian" with its curated, high-quality library performed better than both the smartest AI with internet access and the average human expert.
Why Did the Librarian Win?
The paper explains this using three main ideas:
- Quality Over Quantity: The internet is like a giant bazaar. It has everything, but it's messy. You might find a great recipe next to a fake one. The Librarian's library only has the "gold standard" recipes. When the test asked about a specific, tricky rule for treating diabetes, the Web Surfer might have found a blog post with the wrong info, while the Librarian pulled the exact, up-to-date guideline from the top shelf.
- No Hallucinations: AI models that search the web sometimes "hallucinate"—they make up facts that sound real but aren't. Because the Librarian is locked inside its curated library, it can't make things up. It can only say what is written in its trusted books.
- The "Receipt" System: This is the most important part for doctors. When the Librarian gives an answer, it also hands you a receipt (a citation) showing exactly which page in which book it used.
- Analogy: Imagine a Web Surfer says, "I think the answer is X," but they can't show you where they read it. The Librarian says, "The answer is X, and here is the exact page in the 2024 Diabetes Guidelines that proves it."
- The paper found that 100% of the Librarian's receipts were accurate. This is crucial because doctors can't trust a "black box" AI; they need to verify the proof.
The "Hard Mode" Test
The researchers looked at the 30 hardest questions on the test—the ones that even human doctors got wrong more than half the time.
- The Web Surfer got about 53% of these right.
- The Librarian got 77% of these right.
This suggests that the Librarian's method is especially good at handling the most confusing, complex medical cases where a simple internet search isn't enough.
What Does This Mean for the Future?
The paper argues that for medical AI to be safe and useful, we shouldn't just make the AI "smarter" or give it more internet access. Instead, we should build better, curated libraries for it to learn from.
Think of it like this: If you want to build a self-driving car, you don't just give it a camera and tell it to "watch the road." You program it with a strict, verified map of traffic laws. Similarly, for medical AI, we need to feed it a strict, verified map of medical guidelines.
In Summary:
This study shows that for specialized medical tasks, a focused, verified, and traceable AI system is safer and more accurate than a general, internet-connected AI system. It proves that in medicine, having the right information is more important than having all the information.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.