The strength of clinical evidence is recoverable from language model representations but not from their stated grades
Although large language models encode a recoverable signal of clinical evidence strength within their internal representations, they consistently fail to explicitly state this information, rendering their declared evidence grades unreliable despite the latent signal's existence.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a group of very smart, well-read librarians (the AI models). Their job is to tell you not just what a medical fact is, but how sure we are about that fact. Is it a rock-solid truth backed by thousands of studies, or is it just a hunch from one doctor?
This paper asked a simple question: Do these librarians actually know how sure they are, even if they don't say it out loud?
Here is what the researchers found, broken down into simple concepts:
1. The "Secret Radio" Inside the Librarian
The researchers discovered that every single AI model they tested has a "secret radio" signal inside its brain (its internal computer code). If you know how to tune into this radio, you can hear exactly how strong the evidence is for any claim the AI makes.
- The Catch: The librarians are terrible at speaking this language. When you ask them directly, "How strong is the evidence for this?" they guess randomly. They might say "High confidence" for a weak hunch and "Low confidence" for a rock-solid fact.
- The Analogy: Imagine a weather forecaster who has a perfect, high-tech radar inside their head that knows exactly when a storm is coming. But when they talk to you on the radio, they just guess "Maybe sunny, maybe rainy" with no pattern. The truth is inside them, but they aren't saying it.
2. Bigger Isn't Better (Surprisingly)
Usually, we think bigger, more expensive AI models are smarter. The researchers expected the biggest models to have the clearest "secret radio" signal.
- The Finding: The opposite happened. The smallest models were actually the best at having this signal. The bigger models and the ones specifically trained to "reason" were actually worse at having a readable signal.
- The Analogy: It's like having a small, simple walkie-talkie that gets a crystal-clear signal, versus a massive, high-tech satellite phone that is so complicated the signal gets all scrambled up. The more complex the machine, the harder it is to "read" its internal confidence.
3. The Signal is Mostly Just "Word Choices"
The researchers dug deeper to see what this "secret radio" was actually listening to. They found that the signal wasn't a deep, philosophical understanding of "evidence."
- The Finding: The AI was mostly picking up on the words used in the sentence. Certain phrases (like "studies show" or "expert opinion") act like flags that tell the AI, "Hey, this is a strong claim" or "Hey, this is weak."
- The Analogy: It's like a dog that knows the word "Walk" means "Go outside" and "Bath" means "Stay inside." The dog isn't thinking about the concept of walking or bathing; it's just reacting to the specific sound of the word. The AI is reacting to the vocabulary, not necessarily the deep truth behind the words.
4. The "Truth" vs. "Evidence" Mix-Up
In the real world, a claim can be "strongly supported" but still turn out to be false later (like a medical practice that was popular for years before being proven wrong).
- The Finding: The researchers built a special test to separate "Is this true?" from "Is this well-supported?" They found that the AI's internal signal for "how supported it is" is actually different from its signal for "is it true?"
- The Analogy: Think of a courtroom. The "Evidence Signal" is like the jury's confidence in the quality of the testimony. The "Truth Signal" is the actual verdict of "Guilty" or "Not Guilty." The AI can tell you the testimony was shaky (low evidence) even if the defendant actually did the crime (true). These are two different things, and the AI keeps them separate in its brain, even if it doesn't tell you.
5. The Practical Takeaway
Because the AI won't tell you how strong its evidence is, you can't trust its spoken words. However, because the signal is there inside the machine, a separate, simple computer program can "listen" to the AI's brain and flag weak claims for a human to check.
- The Warning: You cannot trust the AI to grade its own homework. But you can build a simple tool that reads the AI's "internal notes" to catch the mistakes before they reach a doctor or a patient.
Summary
The paper concludes that current medical AI models are like silent experts. They know the difference between a solid fact and a weak guess, but they refuse to say it out loud. They are also mostly just reacting to specific words rather than understanding the deep logic of evidence. Until we can make them speak up or build tools to read their "silent thoughts," we shouldn't trust the confidence levels they claim to have.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.