← Latest papers
💬 NLP

Self-Reported Confidence of Large Language Models in Gastroenterology: Analysis of Commercial, Open-Source, and Quantized Models

This study evaluates the self-reported confidence of various large language models on 300 gastroenterology board-style questions, revealing that while top-performing models show improved accuracy, all models consistently exhibit overconfidence, highlighting a critical barrier to their safe deployment in healthcare.

Original authors: Nariman Naderi, Seyed Amir Ahmad Safavi-Naini, Thomas Savage, Zahra Atf, Peter Lewis, Girish Nadkarni, Ali Soroush

Published 2026-04-10
📖 4 min read☕ Coffee break read

Original authors: Nariman Naderi, Seyed Amir Ahmad Safavi-Naini, Thomas Savage, Zahra Atf, Peter Lewis, Girish Nadkarni, Ali Soroush

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a room full of very smart, very confident robots. These robots are like super-librarians who have read almost every book in the world. You ask them a series of tricky medical questions about the stomach and intestines (gastroenterology).

Here is the twist: Before they give you their answer, you ask them, "On a scale of 0 to 10, how sure are you that you are right?"

This study is like a reality check for these robots. The researchers wanted to see if these AI "librarians" actually know when they are guessing and when they are sure, or if they are just bluffing with a straight face.

The Setup: The Medical Quiz Show

The researchers used a real, difficult medical exam from the American College of Gastroenterology. Think of this as the "Finals" of a medical school test. It has 300 questions that are hard enough to stump even real doctors sometimes.

They tested 48 different AI models. Some were the famous, expensive ones you might have heard of (like GPT-4 or Claude), and some were open-source ones that anyone can download. They asked the robots to answer the questions and rate their own confidence.

The Big Problem: The Overconfident Know-It-Alls

Here is what they found, and it's a bit scary: The robots are terrible at knowing when they are wrong.

Imagine a student taking a test. If they get a question right, they might say, "I'm pretty sure about this!" If they get it wrong, a smart student might say, "Hmm, I'm not sure, maybe I guessed."

But these AI robots are more like a student who gets a question wrong, looks you in the eye, and says, "I am 100% certain I am right," even when they are completely wrong.

  • The "Confidence Gap": The robots were consistently more confident than they were accurate. For example, one of the best robots got about 81% of the answers right, but it claimed to be 91% sure. It was overconfident.
  • The "Blind Spot": Even when the robots got the answer wrong, their confidence score was often just as high as when they got it right. It's like a GPS that confidently tells you to drive off a cliff, even when you are clearly on the wrong road.

The Results: Who Did Best?

The researchers gave the robots grades based on two things:

  1. Accuracy: Did they get the answer right?
  2. Calibration: Did their confidence match their accuracy? (Did they know what they didn't know?)
  • The Winners: The newest, most advanced models (like GPT-o1 and Claude 3.5) were the "least bad." They were better at distinguishing between what they knew and what they didn't, but they were still far from perfect.
  • The Losers: Older or smaller models were often wildly overconfident. They would guess with 90% certainty and be wrong half the time.

Why Does This Matter?

Think about using a robot doctor. If you have a stomach ache, you ask the AI, "What do I have?"

  • If the AI says, "I think it's indigestion, but I'm only 40% sure," you might go see a real human doctor to be safe.
  • But if the AI says, "I am 95% sure you have a stomach ulcer," you might trust it and skip the real doctor.

If the AI is wrong (which happens often in medicine), and it is confidently wrong, that is dangerous. It's like a weatherman confidently predicting a sunny day while a hurricane is forming outside.

The Conclusion: We Can't Trust the "I'm Sure" Button Yet

The study concludes that while these AI models are getting smarter at answering questions, they are not yet smart enough to know when they are guessing.

They are like a student who has memorized the textbook but hasn't learned how to think critically about what they don't know. Until we can teach these robots to say, "I don't know" or "I'm not sure" when they are actually unsure, we can't fully trust them to make life-or-death medical decisions on their own.

In short: The robots are very confident, but they are often confidently wrong. We need to teach them the humility to admit when they are guessing before we let them run the hospital.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →