← Latest papers
💬 NLP

Benchmarking Multi-turn Medical Diagnosis: Hold, Lure, and Self-Correction

This paper introduces MINT, a multi-turn medical diagnosis benchmark that reveals how large language models often rush to premature conclusions despite latent self-correction abilities, and demonstrates that strategically delaying diagnostic questions and evidence presentation can significantly improve diagnostic accuracy.

Original authors: Jinrui Fang, Runhan Chen, Xu Yang, Jian Yu, Jiawei Xu, Ashwin Vinod, Wenqi Shi, Tianlong Chen, Heng Ji, ChengXiang Zhai, Ying Ding, Yuji Zhang

Published 2026-04-07
📖 4 min read☕ Coffee break read

Original authors: Jinrui Fang, Runhan Chen, Xu Yang, Jian Yu, Jiawei Xu, Ashwin Vinod, Wenqi Shi, Tianlong Chen, Heng Ji, ChengXiang Zhai, Ying Ding, Yuji Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery. In a perfect world, you would get the entire case file—the witness statements, the crime scene photos, the lab reports, and the suspect's alibi—all at once. You'd read everything, think it over, and then give your verdict.

Large Language Models (LLMs) are like super-smart detectives who are incredibly good at solving mysteries when they get the whole file at once. But in the real world, detectives don't get all the evidence at once. They get clues one by one: first the witness says, "I saw a red car," then later, "Oh, and the driver was wearing a hat," and finally, "Here is the lab report on the paint chip."

This paper asks a simple but crucial question: Can our AI detectives solve the mystery correctly when the clues are handed to them one by one, or do they get too eager and guess too soon?

The researchers built a new testing ground called MINT (Medical Incremental N-Turn Benchmark) to watch how AI behaves in this "slow reveal" scenario. They discovered three funny but dangerous habits that AI detectives have:

1. The "Impatient Detective" (Hold vs. Rush)

The Problem: When the AI is asked to solve a medical case, it often acts like a child who can't wait for the teacher to finish reading the question. Even if the AI is told, "Wait until you have all the clues before you guess," it often blurts out an answer after seeing just the first two or three pieces of information.

The Analogy: Imagine a game show where the host gives you a riddle. The AI is like a contestant who shouts out an answer after hearing only the first word of the riddle, just because it wants to be right.

  • The Finding: Over 55% of the time, the AI guesses within the first two turns. If you force the AI to wait until the very end to give its answer (like making the detective wait until all evidence is in), its accuracy jumps back up to near-perfect levels. The AI knows the answer; it just can't stop itself from guessing early.

2. The "Second-Guessing Genius" (Self-Correction)

The Problem: When the AI guesses early, it often gets it wrong. But here's the cool part: as more clues come in, the AI is actually really good at changing its mind and fixing its mistake.

The Analogy: Think of the AI as a student taking a test. They might circle "B" on a question immediately (getting it wrong). But as they read the next paragraph of the textbook, they realize, "Oh wait, it's actually 'C'!" and they change their answer.

  • The Finding: The AI is 10.6 times more likely to change a wrong answer to a right one than to change a right answer to a wrong one. It has a hidden superpower of self-correction, but this power is useless if it's too impatient to wait for the new clues to arrive.

3. The "Shiny Object Trap" (Strong Lures)

The Problem: Some pieces of evidence are just too shiny. In medicine, things like Lab Results (blood tests, X-rays) are very specific and look like "the answer." When the AI sees a lab result early in the conversation, it gets distracted and immediately jumps to a conclusion, even if it hasn't heard the patient's story yet.

The Analogy: Imagine you are looking for a lost dog. You get a clue: "The dog was seen near a park." You might guess the dog is in the park. But then, suddenly, someone hands you a photo of the dog's collar with a specific name tag. Even if you haven't heard the rest of the story, your brain screams, "I know where the dog is!" and you stop looking.

  • The Finding: Lab results act as a "lure." When they appear early, the AI stops thinking and starts guessing. This causes a massive drop in accuracy, especially for diseases where the lab test isn't actually the most important clue.

The Big Takeaway

The paper concludes that the AI isn't "bad" at medicine. It's actually quite smart. The problem is timing.

  • The Old Way: We ask the AI to solve the case immediately, and it rushes, makes mistakes, and can't recover.
  • The New Advice: We should teach the AI to wait. If we design systems that tell the AI, "Don't give me your final diagnosis until you've seen the last piece of evidence," the AI becomes incredibly reliable.

In short: The AI is a brilliant detective who just needs to be told, "Sit down, drink your coffee, and read the whole file before you speak." If we do that, it can save lives.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →