← Latest papers
🧬 biology

Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries

A blinded evaluation by 149 physicians across 30 specialties demonstrates that a specialized clinical AI tool significantly outperforms three frontier general-purpose large language models on real-world point-of-care queries, highlighting the critical importance of using expert judges and authentic clinical data for evaluating medical AI.

Original authors: Jean Feng, Vishal Patel, Patrick Heagerty, Yifan Mai, Venkatesh Sivaraman, Patrick Vossler, Jialin Ouyang, Anupam B. Jena

Published 2026-06-30
📖 4 min read☕ Coffee break read

Original authors: Jean Feng, Vishal Patel, Patrick Heagerty, Yifan Mai, Venkatesh Sivaraman, Patrick Vossler, Jialin Ouyang, Anupam B. Jena

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are a chef trying to decide which recipe book is best for your restaurant. You have two options:

  1. The "Generalist" Books: Massive encyclopedias that know a little bit about everything (cooking, history, science, and even how to fix a car).
  2. The "Specialist" Book: A cookbook written specifically for your type of cuisine, filled with tips from master chefs who only cook that specific food.

Usually, when people test these books, they ask them questions like, "What is the capital of France?" or "How do you boil an egg?" These are easy, made-up questions. But in the real world, a chef doesn't ask, "How do you boil an egg?" They ask, "My soup is too salty, but I can't add more water because the broth is already perfect; what do I do?"

This paper is a massive, blind taste test to see which "recipe book" (AI tool) actually helps doctors solve real, messy problems they face every day.

The Setup: A Blind Taste Test

The researchers gathered 620 real questions that doctors actually asked a specialized medical AI tool called OpenEvidence (OE) while treating patients. These weren't made-up exam questions; they were the actual things doctors needed to know right now.

They took these questions and fed them to four different AI "chefs":

  • Three Generalists: The latest, most powerful "do-everything" AIs (GPT-5.5, Gemini 3.1 Pro, and Claude Opus 4.8).
  • One Specialist: OpenEvidence (OE), a tool built specifically for doctors.

Then, they hired 149 real doctors from 36 different states to act as the judges. To make it fair, they matched the judges to the questions. If the question was about heart disease, a cardiologist graded the answers. If it was about skin issues, a dermatologist graded them. The doctors didn't know which AI wrote which answer (it was "blinded"), just like a judge in a cooking competition doesn't know the chef's name until the end.

The Results: The Specialist Wins

The doctors rated the answers on five things:

  1. Accuracy: Was the fact right?
  2. Usefulness: Could I actually use this to help a patient?
  3. Source Quality: Did it cite good, trustworthy books or journals?
  4. Verifiability: Could I easily check if it was true?
  5. Completeness: Did it answer the whole question?

The Outcome: The specialized tool (OpenEvidence) won every single category. It beat the generalist giants by a huge margin.

  • In the "Accuracy" category, the specialist was preferred over the generalists by about 25% to 39% more often.
  • In "Source Quality," the specialist won by nearly 40%.

The generalist AIs (the "do-everything" models) didn't do terribly, but they were consistently outperformed by the tool built specifically for the job.

The "Robot Judge" Experiment

The researchers also tried something interesting: they asked the AI models to grade each other's answers (a "Robot Judge").

  • The Problem: The Robot Judges were often overconfident and biased. For example, the GPT model tended to give itself high scores, even when human doctors thought it was wrong.
  • The Lesson: While Robot Judges agreed on who was the best overall, they disagreed with humans on the details. You can't just let an AI grade another AI without checking it against real human experts.

Why This Matters (According to the Paper)

The paper makes two main points:

  1. Real Tests Need Real Questions: If you only test AI on made-up exam questions, you don't know how it will perform in the real world. You need to test it on the messy, specific questions doctors actually ask.
  2. Specialization Wins: Just because a general AI is smart doesn't mean it's the best tool for a specific job. The fact that the specialized tool won doesn't mean the generalists are useless; it means that customizing an AI for a specific field (like medicine) makes it much better at that specific job.

What the Paper Doesn't Say

  • It does not say these tools should replace doctors.
  • It does not say the generalist AIs are "bad" at everything; they just weren't as good as the specialist for these specific medical questions.
  • It does not claim this is the final word on AI, as the models change quickly.

In short: When you need a doctor to help with a specific medical problem, a tool built just for doctors works better than a general "smart" AI that tries to do everything. And to find out which tool is best, you have to test them with real doctors and real questions, not just fake exams.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →