← Latest papers
💬 NLP

Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

This paper introduces PhysAssistBench, a novel benchmark constructed from real MIMIC-IV cases that evaluates large language models on their ability to coordinate clinical knowledge, patient communication, and EHR tool use in interactive doctor-patient scenarios, revealing that current models remain unreliable for such integrated physician assistance despite isolated improvements in individual capabilities.

Original authors: Tianming Du, Peijie Yu, Sihan Shang, Danli Shi, My Linh Nguyen, Shengbo Gao, Guangyuan Li, Yinghong Yu, Yan Jiang, Qianlong Zhao, Behzad Bozorgtabar, Shaoxiong Ji, Jiazhen Pan, Daniel Rueckert, Jianch
Published 2026-06-19
📖 5 min read🧠 Deep dive

Original authors: Tianming Du, Peijie Yu, Sihan Shang, Danli Shi, My Linh Nguyen, Shengbo Gao, Guangyuan Li, Yinghong Yu, Yan Jiang, Qianlong Zhao, Behzad Bozorgtabar, Shaoxiong Ji, Jiazhen Pan, Daniel Rueckert, Jiancheng Yang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: The "Super-Intern" Test

Imagine a hospital where doctors are overwhelmed. They want to hire a "Super-Intern" (an AI) to help them. This intern needs to do three things at once:

  1. Read the patient's file (Electronic Health Records or EHR) instantly.
  2. Talk to the patient to get the full story.
  3. Listen to the Doctor, who is busy and often speaks in shorthand.

The paper argues that while AI is great at taking medical exams (like a student memorizing a textbook), we don't actually know if it can handle the messy, real-life job of being a doctor's assistant. To find out, the authors built a new, very difficult test called PhysAssistBench.

The Problem: The "Textbook" vs. The "Real World"

Think of current AI tests like a driving exam where you only have to park a car in an empty lot with perfect cones. The AI passes with flying colors.

But real life isn't an empty lot. It's rush hour traffic.

  • The Doctor: Instead of saying, "Please check the blood pressure," the doctor might just say, "How's the pressure?" or even just "Pressure?" (This is called an implicit query).
  • The Patient: Instead of saying, "I have high blood pressure," the patient might say, "My head feels like a balloon and my socks leave deep marks on my ankles." (This is ambiguous communication).
  • The System: The hospital computer requires you to click specific buttons in a specific order to get the data.

The paper says current AI models fail when you mix these three things together. They get lost in the traffic.

The Solution: A "Video Game" Hospital

To test the AI properly, the researchers built a realistic simulation using real, anonymized patient data from a database called MIMIC-IV.

They didn't just write questions; they created a video game environment with three characters:

  1. The Busy Doctor: An AI that asks short, vague questions based on real medical cases.
  2. The "Agentic" Patient: A computer character that acts like a real human. It has a medical file, but it also has a personality. It might forget to mention a symptom or describe it in slang. It answers questions based only on its real medical history, not made-up stories.
  3. The Hospital Computer: A strict system that only gives data if you ask for it using the exact right digital "keys" (tools).

The AI being tested has to play the role of the Assistant. It has to listen to the Doctor, figure out what they actually mean, ask the Patient the right questions, check the Computer for the facts, and then give the Doctor a clear answer.

The Test: Four Rounds of Chaos

The test consists of 324 different "scenarios" (like different patient cases). Each scenario has four rounds:

  • Round 1: The Doctor asks for a specific fact (e.g., "What's the latest blood test?").
  • Round 2: The Doctor asks for more info, but uses shorthand (e.g., "And the meds?").
  • Round 3: The Doctor asks for a recommendation based on everything so far (e.g., "Given all this, what do we do?").
  • Round 4: The Doctor asks the AI to write a new prescription or update the file.

The AI has to get all four rounds right to pass the whole scenario. If it messes up just one turn, the whole session fails.

What Happened? The "Super-Intern" Stumbles

The researchers tested 14 of the smartest AI models available (including big names like GPT-5, Claude, and Gemini).

The Results:

  • The Good News: The AI is great at simple tasks. If the Doctor asks, "What is the blood pressure?" and the AI just looks it up, it gets it right 80%+ of the time.
  • The Bad News: When the test gets complex, the AI struggles badly.
    • The "Shorthand" Problem: When the Doctor uses vague language (like "Check the meds"), the AI often gets confused about which meds or what to check.
    • The "Patient" Problem: When the AI has to talk to the "Patient" to get missing info, its performance drops significantly. It's much better at reading a computer file than having a conversation.
    • The "All-or-Nothing" Problem: Even the best models only passed about 8% to 23% of the entire 4-round scenarios perfectly. This means that in a real hospital, the AI would likely make a mistake in a multi-step conversation more often than it would get it right.

The Conclusion

The paper concludes that AI is not yet ready to be a reliable "co-pilot" for doctors in a real hospital.

The Analogy:
Imagine you are teaching a robot to be a chef.

  • Old Tests: You asked the robot, "Can you chop an onion?" It passed.
  • This New Test: You put the robot in a busy kitchen. The Head Chef yells, "Fix the soup!" The robot has to taste the soup, ask the customer what they want, check the pantry for ingredients, and then cook it.
  • The Result: The robot keeps burning the soup or forgetting to ask the customer. It knows how to chop onions, but it doesn't know how to run the kitchen.

The authors say the biggest hurdle isn't that the AI doesn't know enough medicine; it's that it can't coordinate listening, talking, and using tools all at the same time without getting confused. They have released this test to the public so other researchers can try to fix these specific problems.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →