← Latest papers
🧬 biology

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

This paper demonstrates that DR. INFO, an agentic RAG-based medical assistant, significantly outperforms leading frontier LLMs and competing clinical AI systems on the HealthBench Hard benchmark, validating the effectiveness of rubric-driven, open-ended evaluations for assessing high-stakes clinical reasoning and communication.

Original authors: Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam

Published 2026-07-28
📖 3 min read☕ Coffee break read

Original authors: Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are trying to teach a super-smart robot how to be a doctor. For a long time, we tested these robots using multiple-choice quizzes, like the ones you might see in a high school biology exam. You'd ask, "What is the capital of France?" or "Which medicine treats a headache?" and the robot would pick A, B, or C. If it got the answer right, we assumed it was ready to help real people. But here's the catch: real life isn't a multiple-choice quiz. Real life is messy, confusing, and full of missing details. A patient might be scared, forget to mention a symptom, or ask a question in a way that doesn't fit a neat box. If a robot just memorizes facts but can't listen, ask the right follow-up questions, or admit when it's unsure, it could make dangerous mistakes. This is where a new kind of test comes in, designed not to see if the robot knows the textbook answers, but to see how it behaves in a real, high-pressure conversation.

This paper is about a team at a company called Synduct who built a medical robot assistant named DR. INFO. They wanted to see if their robot was actually ready for the real world, so they put it through a new, tougher test called HealthBench. Unlike the old quizzes, HealthBench is like a giant role-playing game with 1,000 difficult scenarios. In these scenarios, a human plays a patient (or a doctor) and asks a tricky question, and the robot has to respond. A team of real doctors then grades the robot's answer based on a detailed checklist (called a rubric) that looks at things like: Did the robot tell the truth? Did it ask for more information when things were unclear? Did it stay calm and clear? Did it follow instructions?

The results were a huge surprise. When DR. INFO took this tough test, it scored a 0.68 out of 1.0. To put that in perspective, the paper compared DR. INFO to the very best, most famous AI models in the world right now, including OpenAI's GPT-5 family. The GPT-5 models scored much lower, with the best version getting only 0.46. DR. INFO didn't just beat them; it substantially outperformed them, achieving a score that represents a 48% relative increase over the top competitor. It was even better at specific skills like knowing when to ask for more context (scoring 0.62 vs. 0.16 for another top model) and giving complete answers. The team also tested DR. INFO against other medical robots designed to do similar jobs, and DR. INFO won those races too, scoring 0.72 compared to their rivals' scores of around 0.49.

However, the authors are careful not to say the robot is perfect or that the problem is "solved." They found that while DR. INFO is great at following instructions and being accurate, it still has some room to grow in how well it understands the full context of a conversation and how complete its answers are. The paper suggests that this new way of testing—looking at behavior rather than just facts—is the key to building safe, trustworthy medical AI. It's like realizing that to be a good driver, you don't just need to know the rules of the road; you need to know how to react when a squirrel jumps in front of your car. DR. INFO seems to be learning how to handle the squirrels better than the other robots, but the journey to becoming a fully autonomous medical assistant is still ongoing.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →