← Latest papers
🤖 AI

LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap

This paper argues that large language models for medical consultation are currently evaluated too late in the process, demonstrating that explicit instructions can improve structured documentation but fail to reliably prevent premature self-care advice or ensure the elicitation of critical facts during the initial "preformulation" phase of a vague patient concern.

Original authors: Yining Hua, Cyrus Ayubcha, Hongbin Na, Levi Lian, Alon Gorenshtein, Yiftach Barash, Eyal Klang

Published 2026-08-19
📖 5 min read🧠 Deep dive

Original authors: Yining Hua, Cyrus Ayubcha, Hongbin Na, Levi Lian, Alon Gorenshtein, Yiftach Barash, Eyal Klang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a person feeling unwell, sitting alone with a smartphone, typing a few hurried words into a chat window. They might misspell a symptom, downplay their pain, or describe a vague feeling of sickness without knowing the medical name for what is wrong. In the real world, a doctor's first job is not to guess the disease immediately, but to ask the right questions to turn that messy, confusing story into a clear medical picture. This process of turning a vague worry into a specific, actionable problem is the foundation of safe care. If a doctor skips this step and jumps straight to advice, they might miss a critical detail that changes everything. Now, imagine a computer program designed to act like a doctor. For years, researchers have tested these programs by giving them clear, well-written medical questions and grading how well they answer. But this method leaves out the most dangerous moment: the very first exchange, when the patient's concern is still raw, unformed, and potentially misleading.

A team of researchers from Harvard and other institutions recently decided to look at this missing moment. They asked a simple but vital question: when a person types a messy, incomplete message into a medical chatbot, does the computer stop to ask clarifying questions, or does it immediately start giving advice based on the wrong assumption? To find out, they created four realistic scenarios where a patient described a symptom in a way that sounded harmless but hid a serious danger. In one case, a person wrote that they were "throwing up" and thought it was food poisoning, but they were actually suffering from a life-threatening complication of diabetes. In another, someone described leg pain after a long trip, which sounded like a muscle strain but was actually a blood clot. The researchers tested three different large language models—the powerful computer systems that power popular chatbots—under two conditions. First, they let the models run naturally, as they would for a regular user. Second, they gave the models a specific set of instructions to follow before answering, telling them to ask questions first and check for safety risks.

The results revealed a significant gap in how these systems behave. When left to their own devices, the models often made a critical mistake: they treated the patient's vague description as a complete story and immediately offered home-care advice. In nine out of twelve test cases, the computer suggested things like drinking water, eating bland food, or resting before it had even asked a single follow-up question. This happened even when the patient's initial message was misspelled and lacked crucial details. The models were so eager to be helpful that they filled in the blanks with assumptions, potentially steering a sick person away from the emergency room and toward a dangerous wait-and-see approach. The researchers found that this happened even when the models eventually got the right answer later in the conversation; the damage was done in that first moment when the patient was given false reassurance.

However, the study also showed that this behavior is not fixed in the code; it can be changed with a simple shift in how the task is framed. When the researchers gave the models a short instruction to ask key questions before offering advice, the behavior flipped completely. In the instructed group, none of the models gave home-care advice before asking questions. Instead, they paused to inquire about age, pain severity, and medical history. They also became much better at spotting risky plans, such as a patient saying they intended to "sleep it off" or drink alcohol, and they explicitly warned against these actions. Furthermore, the instructed models began to create a clear summary of the situation at the end of the chat, a "handoff" that a real doctor could read to understand what had happened. This summary was missing entirely in the unguided tests.

Despite these improvements, the researchers discovered that a simple instruction was not a perfect fix. Even when told to ask questions, the models sometimes failed to dig deep enough for the most critical facts. In the case of the vomiting patient with diabetes, the models asked questions, but they did not specifically ask about insulin or diabetes until the patient volunteered that information later in the chat. This suggests that while the models can be trained to change their order of operations, they do not yet reliably know which specific questions are the most important to ask when a patient is vague. The study concludes that the current way we test medical chatbots is too late. By only grading the final answer, we miss the dangerous gap between a patient's first confused message and the moment a doctor—or a computer—forms a clear picture of the illness. The researchers argue that for these tools to be safe, we must evaluate them on their ability to listen and question first, not just on their ability to diagnose and answer. The safety of the patient depends on what happens in those first few seconds of conversation, long before the final verdict is delivered.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →