← Latest papers
💻 computer science

Possible or Definite? A Benchmark for Evaluating Diagnostic Uncertainty Preservation in Clinical Text

This paper introduces a benchmark of 1,200 clinical documents to evaluate how well large language models preserve diagnostic uncertainty, revealing that current models frequently fail to maintain these critical nuances, thereby posing risks for safe clinical deployment.

Original authors: Hongbo Du, Zixin Lu, Jiaming Qu

Published 2026-06-19
📖 5 min read🧠 Deep dive

Original authors: Hongbo Du, Zixin Lu, Jiaming Qu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: The "Maybe" vs. "Definitely" Problem

Imagine you are a doctor writing a report about a patient. You aren't 100% sure if they have pneumonia, so you write: "Possible pneumonia."

In the medical world, that word "possible" is a safety valve. It tells the next doctor, "We need to keep looking; don't start heavy treatment yet." If you change that to just "Pneumonia," the meaning shifts entirely. Now, the next doctor might think, "Okay, it's confirmed," and start treatment that isn't needed or skip important tests.

This paper asks a simple but scary question: If we ask a smart AI (a Large Language Model) to rewrite or summarize these medical notes, does it keep the "maybe" intact, or does it accidentally turn "maybe" into "definitely"?

The Experiment: Building a "Stress Test" for AI

The researchers didn't just guess; they built a massive "stress test" to see how the AI behaves.

  1. The Training Ground: They gathered 1,200 real medical documents (like discharge summaries and lab reports) from hospitals and cancer databases.
  2. The Map: They marked over 9,000 specific spots in these texts where a doctor expressed uncertainty. They created a 5-level scale, ranging from:
    • Level 1 (Certain Absent): "No, the patient definitely does not have this."
    • Level 2 (Probable): "It's likely, but not 100%."
    • Level 3 (Possible): "It could be, but we aren't sure."
    • Level 4 (Indeterminate): "We don't have enough info to tell."
    • Level 5 (Non-Asserted): "We need to check for this later."
  3. The Test: They fed these documents into three popular AI models (from OpenAI, Google, and Anthropic) and asked them to do two things:
    • Summarize: Write a short version for other doctors.
    • Rewrite: Translate the jargon into plain English for the patient.

They tested the AI in two ways:

  • The "Normal" Way: Just asking it to summarize/rewrite.
  • The "Guarded" Way: Giving it a strict rule: "Do not change the level of certainty. If it says 'maybe,' keep it as 'maybe'."

The Results: The AI is a "Confident" Over-Editor

The results showed that the AI has a serious habit of being too confident.

1. The "Confidence Trap"
When the AI rewrote the text, it often deleted the "maybe" words entirely.

  • The Analogy: Imagine a weather forecaster saying, "There is a possible chance of rain." If you ask an AI to summarize that for a picnic planner, it might say, "It will rain." It didn't lie about the rain existing, but it removed the uncertainty, making the forecast sound like a guarantee.
  • The Stat: Even when the AI kept the medical fact (e.g., "pneumonia"), it only kept the correct level of uncertainty less than 50% of the time.
  • The Worst Offense: About 40% of the time, the AI took a "possible" diagnosis and turned it into a "definite" one. This is the most dangerous kind of error because it sounds authoritative but is actually a guess.

2. The "Patient" vs. "Doctor" Gap
The AI was worse at rewriting text for patients than for doctors.

  • The Analogy: When talking to doctors, the AI acted like a careful editor, keeping the nuances. When talking to patients, it acted like a storyteller trying to simplify the plot, often cutting out the "plot holes" (uncertainties) to make the story flow better.
  • The Result: The AI dropped more medical facts and changed more uncertainties when trying to make the text "patient-friendly."

3. The "Magic Prompt" Didn't Work
The researchers tried to fix this by giving the AI a strict instruction: "Please preserve the uncertainty!"

  • The Result: It helped a little bit (improving accuracy by about 10–20%), but it didn't fix the problem. The AI still turned "possible" into "definite" far too often. It's like telling a student, "Don't guess on the test," and they still guess anyway.

4. The "Ranking" Test
In a second test, they asked the AI to look at a list of sentences and rank them from "Most Certain" to "Least Certain."

  • The Result: The AI was okay at telling the difference between "Definitely No" and "Definitely Yes." But it got very confused when trying to tell the difference between "Probable" and "Possible." It struggled to see the fine lines between similar levels of doubt.

The Takeaway

The paper concludes that while these AI models are great at sounding fluent and making sentences that look grammatically perfect, they are currently bad at preserving the "shades of gray" in medical language.

They tend to turn "maybe" into "yes." In a hospital, that isn't just a typo; it's a change in the medical meaning that could lead to the wrong tests or treatments. The authors warn that we cannot just trust these AI tools to summarize medical notes until we can teach them to respect the word "possible" just as much as the word "definite."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →