← Latest papers
🤖 AI

Large language models improve physician accuracy but lead to false reliance

While the retrieval-augmented LLM system CORA significantly improved physician diagnostic accuracy, the study reveals a critical safety risk where source-linked citations disproportionately increased reliance on both correct and incorrect advice, thereby undermining physicians' ability to resist erroneous recommendations.

Original authors: Tirtha Chanda, Christoph Wies, Franziska Schramm, Carina Nogueira Garcia, Nicolas B. Merl, Martin J. Hetz, Jochen S. Utikal, Phillip Tschandl, Cristian Navarrete-Dechent, Alexander Thiem, Jakob N. Kat
Published 2026-08-04
📖 5 min read🧠 Deep dive

Original authors: Tirtha Chanda, Christoph Wies, Franziska Schramm, Carina Nogueira Garcia, Nicolas B. Merl, Martin J. Hetz, Jochen S. Utikal, Phillip Tschandl, Cristian Navarrete-Dechent, Alexander Thiem, Jakob N. Kather, Consortium, Titus J. Brinker

Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a tricky case. You have a brilliant, super-fast assistant who can read millions of books in a second and suggest a solution. This assistant is an Artificial Intelligence, specifically a "Large Language Model" (LLM). These AIs are like encyclopedias that learned to talk by reading the entire internet, but sometimes they get confident about things they don't actually know, making up facts that sound perfect but are wrong. To fix this, scientists invented a trick called "Retrieval-Augmented Generation" (RAG). Think of RAG as giving your AI assistant a library card and a rule: "You can only answer if you can point to the exact page in a book that proves your answer." This way, the AI doesn't just guess; it shows its homework.

But here is the big question that scientists are still debating: Does showing the homework actually help the human detective? Or does it just make the human trust the assistant too much, even when the assistant is wrong? This is the story of a new study that put this idea to the test with real doctors. They wanted to see if giving doctors an AI that cites its sources would make them better at diagnosing skin diseases, or if it would trick them into making dangerous mistakes by making wrong answers look like they were backed by evidence.


The Story of CORA and the "Fake Proof" Trap

In this study, a team of researchers built a special AI assistant named CORA (Citation-Oriented Retrieval Assistant). Imagine CORA as a super-smart research intern who, whenever a doctor asks a question about a skin condition, doesn't just spit out an answer. Instead, CORA goes to a digital library of medical textbooks, guidelines, and case reports, finds the relevant pages, and says, "Here is my answer, and here are the three pages that prove it."

The researchers wanted to see how this worked in the real world. They gathered 46 doctors from 21 different countries and gave them a series of skin disease puzzles. First, the doctors had to solve them on their own. Then, they were shown CORA's answer along with the "proof" (the citations) and asked if they wanted to stick with their original answer or change it.

The Good News: Doctors Got Smarter
The results showed that CORA was a helpful sidekick. When the doctors worked alone, they got the right answer 70.8% of the time. But when they used CORA, their accuracy jumped to 82.6%. The AI helped them fix mistakes, especially on tricky cases or diseases that were rare. It was like having a second pair of eyes that knew the rulebook perfectly.

The Bad News: The "Citation Trap"
However, the study found a sneaky problem. The "proof" CORA showed wasn't always real proof. Sometimes, CORA would give a wrong answer but still show a citation that looked like it supported the answer, even though it didn't really.

Here is where it gets tricky:

  • When CORA was right: If the doctors saw a citation that seemed to support the answer, they were very likely to change their mind and agree with the AI. Their success rate for adopting the correct advice went from 34% (without feeling supported) to 76.9% (when they felt the citation supported it).
  • When CORA was wrong: This is the dangerous part. If CORA gave a wrong answer but showed a citation that looked supportive, the doctors stopped resisting. Normally, doctors are stubborn and stick to their correct answer 92% of the time when the AI is wrong. But when a "supportive" citation was shown, that resistance dropped to just 34.8%.

The Lesson: Trust, but Verify the Proof
The study discovered a "grounding-miscalibration." This is a fancy way of saying that the doctors were fooled by the appearance of evidence. When the AI showed a citation, the doctors assumed the answer must be right, even if the citation didn't actually prove the specific point. It's like a student handing in a test with a textbook page attached. If the page is about the right topic but doesn't actually answer the question, a teacher might still think, "Oh, they did the research," and give them credit. But in medicine, that credit could be a mistake.

The researchers found that while the AI made doctors more accurate overall, it also created a new risk: doctors were more likely to give up their own correct instincts if the AI made a wrong answer look well-supported. The study suggests that for these AI tools to be safe, we can't just look at whether the AI is right or wrong. We have to make sure the "proof" it shows actually proves the answer, not just that it's related to the topic. Otherwise, the very thing meant to help doctors verify the truth might end up tricking them into trusting a lie.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →