← Latest papers
💻 computer science

CARE-Bench: Benchmarking Patient-Facing LLM Triage

The paper introduces CARE-Bench, a source-grounded benchmark for evaluating patient-facing medical LLMs on sequential triage tasks, revealing that while prompting improves performance, models still frequently fail to correctly time care recommendations versus information gathering, highlighting significant safety risks for deployment.

Original authors: Yining Hua, Hongbin Na, Cyrus Ayubcha

Published 2026-08-05
📖 4 min read☕ Coffee break read

Original authors: Yining Hua, Hongbin Na, Cyrus Ayubcha

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are standing in a crowded, noisy waiting room, but instead of people, it's filled with digital assistants—smart, chatty bots that promise to help you figure out what's wrong with your health. You type, "My head hurts," and they instantly reply with a long list of possibilities, from "drink some water" to "call an ambulance." This is the world of patient-facing medical AI, where large language models (LLMs) act as the first line of defense before you ever see a real doctor. The big question isn't just whether these bots know medical facts; it's about triage. Think of triage like a traffic cop at a busy intersection. The cop doesn't need to know how to fix a broken engine; they just need to know: Do you keep driving, pull over to check the tire, or get out and call a tow truck right now? If the cop sends a minor fender-bender to the tow truck, traffic jams (unnecessary hospital visits) happen. If they send a car with a flat tire straight to the highway, disaster strikes. The challenge is teaching these digital cops to know exactly when to ask for more info, when to say "rest up," and when to scream "emergency!"

The paper you're about to hear about, CARE-BENCH, is like a giant, high-stakes driving test for these medical bots. The researchers built a special "simulated road" using 500 real-life medical stories, broken down into small, step-by-step conversations. They didn't just ask the bots, "What's the answer?" Instead, they watched how the bots reacted at every single turn. Did the bot ask the right clarifying questions when the patient was vague? Did it panic and send someone to the ER too early? Or did it stay too calm when a patient was actually in danger? They tested 11 different AI models, some from big tech companies and some open-source, under two conditions: one where the bots got no special instructions (just like a real user typing a question), and one where they got a tiny nudge to "be brief and safe."

Here is the twist: The bots are still getting the traffic cop job wrong. Even the smartest models, when left to their own devices, scored poorly, with a "macro-F1" (a fancy score for how well they balance different types of mistakes) ranging only from 31.2 to 50.4. When the researchers gave them a simple prompt to "be safe," the scores improved for 10 out of 11 models, reaching up to 63.4. But here is the catch: prompting didn't fix the core problem. The bots still made dangerous timing errors. When a patient gave a vague description of symptoms, the correct move was to ask, "Wait, tell me more about that," before giving advice. But the bots often skipped this step entirely. They jumped straight to giving advice—sometimes telling people to go to the ER when they didn't need to, or telling them to stay home when they should have rushed to the hospital.

In fact, when the right move was to ask for more information, the prompted bots only got it right 33.5% of the time. They were so eager to be "helpful" that they forgot to be "careful." The study suggests that simply telling an AI to "be safe" isn't enough to make it a good triage nurse. The error isn't just about knowing medical facts; it's about timing. The bots struggle to realize that they don't have enough information yet to make a decision. They treat a vague symptom like a full diagnosis, leading to a mix of unnecessary panic and dangerous delays. The researchers conclude that before we let these bots loose in the real world to decide who goes to the hospital, we need to test them much more rigorously on when they speak, not just what they say. The current models are like enthusiastic but inexperienced interns who are too eager to give a prescription before they've even finished listening to the patient's story.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →