Active Evidence-Seeking and Diagnostic Reasoning in Large Language Models for Clinical Decision Support
This paper introduces an OSCE-inspired benchmark for active diagnostic inquiry, revealing that large language models experience significant drops in accuracy and evidence quality during multi-turn evidence seeking compared to static full-context evaluations, primarily due to premature diagnostic closure and inefficient questioning.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: The "All-Seeing" vs. The "Detective"
Imagine you are taking a medical exam.
Scenario A (The Old Way): You are handed a thick file folder containing the patient's entire life story, every symptom they've ever had, every blood test result, and every X-ray image. You read the whole thing and write down the diagnosis. This is how most AI models are currently tested. They are great at this "file folder" style test.
Scenario B (The Real World): You walk into a room with a patient. You only know they have a "chest pain." You don't have the file folder yet. You have to ask questions: "Does it hurt when you breathe?" "Do you have a history of smoking?" You have to order specific tests one by one. You have to build the file folder yourself, piece by piece, based on what you ask.
The Paper's Discovery: The researchers built a new test called ROUNDS-Bench to see how AI handles Scenario B. They found a shocking gap: AI models that are geniuses at Scenario A often fail miserably at Scenario B.
When forced to play the role of a detective who has to ask for clues, the AI's accuracy dropped by about 13%, and the quality of the evidence they gathered dropped by nearly 25%.
The Experiment: A "Standardized Patient" Simulator
To test this fairly, the researchers created a digital "Standardized Patient" (like a method actor in a play).
- The Actor: This AI patient knows the full medical story but is programmed to only reveal information if asked the right questions.
- The Rules: If the doctor (the AI model) asks, "Tell me everything," the patient says, "I can't do that." The doctor must ask specific questions like, "Can you describe your pain?" or "Let's check your heart rate."
- The Goal: The AI has to figure out the disease by asking the right questions in the right order, just like a real doctor.
They tested 15 different AI models (including big names like GPT-4o, Gemini, and Qwen) on 468 different medical cases covering everything from heart issues to infections.
What They Found
1. The "Crash" in Performance
When the AI was given the full file folder (Scenario A), it got high scores. But when it had to ask for clues (Scenario B), its performance tanked.
- The Analogy: Imagine a student who gets an 'A' on a multiple-choice test because they can read the whole textbook. But if you put them in a room with a mystery and tell them they can only ask three questions to solve it, they might guess the wrong answer entirely.
- The Result: The AI didn't just get "a little less right"; it often went from being 100% correct to being completely wrong. It stopped being cautious and started making wild guesses based on very little information.
2. The "Magic Trick" of Hallucinated Reasoning
This is the most dangerous finding. Sometimes, the AI would guess the correct disease, but it couldn't explain why it was right based on the clues it actually found.
- The Analogy: Imagine a detective solves a murder mystery. They say, "The butler did it!" (Correct). But when asked for proof, they say, "Because I saw a shadow," even though the shadow was never mentioned in the case file. They just guessed the right answer for the wrong reasons.
- The Risk: In a real hospital, if a doctor trusts an AI that says "It's this disease" but can't point to the specific test result that proves it, that's a safety hazard. The paper calls this "Hallucinated Reasoning."
3. Size Doesn't Always Save You
You might think bigger, smarter AI models would be better detectives.
- The Reality: While bigger models did slightly better than tiny ones, even the biggest, most expensive models still struggled. They were good at reading the whole file, but bad at knowing what to ask next when they didn't have the file.
- The "Distilled" Trap: Some models are trained to "think step-by-step" (reasoning models). The paper found that while these models are great at solving puzzles when all the pieces are on the table, they get confused when they have to go out and find the missing pieces.
4. Some Diseases Are Harder to "Detect"
The AI did okay with things like heart and lung issues (which often have clear, standard symptoms). But it fell apart with Neurological (brain/nerves) and Metabolic (blood chemistry) cases.
- Why? These areas require very specific, nuanced questions (like "Is your reflex weak in the left foot?") or complex math with lab numbers. The AI struggled to know exactly which specific question to ask to get the right answer.
The Bottom Line
The paper argues that we have been praising AI doctors for being great at reading medical files, but we haven't tested them enough on doing the work of a doctor (asking questions and gathering clues).
The Takeaway: Just because an AI can pass a written medical exam doesn't mean it can safely talk to a patient and figure out what's wrong. To make AI safe for real hospitals, we need to stop testing them with "file folders" and start testing them as "detectives" who have to earn their answers.
The researchers created this new test (ROUNDS-Bench) to help developers fix these "detective skills" so that future medical AI doesn't just guess the right answer, but actually knows why it's right.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.