When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs
This paper proposes a decision-theoretic framework to validate the consistency between Large Language Models' elicited probability beliefs and their decisions, revealing that while even the strongest models exhibit minor discrepancies between their stated beliefs and actions, these beliefs are not perfect summaries of the information driving their choices.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a medical consultant who is an expert at diagnosing patients. You ask them two things about a specific patient:
- The Guess: "What is the percentage chance this patient has a specific disease?"
- The Action: "Based on what you know, what should we do? Treat them, send them home, or get more tests?"
This paper asks a simple but tricky question: Does the consultant's "Guess" actually match their "Action"?
Sometimes, people (or AI) say one thing but do another. Maybe the consultant says, "I'm only 60% sure it's the disease," but then they act as if they are 99% sure by immediately prescribing a risky surgery. Or maybe they act cautiously, even though they claim to be very confident.
The researchers in this paper built a "truth detector" to see if Large Language Models (LLMs)—the smart AI chatbots—are being consistent. They didn't just check if the AI's math was right; they checked if the AI's words (its stated beliefs) actually explained its choices.
The "Black Box" Detective Work
The researchers treat the AI like a "black box." They can't look inside its brain to see how it thinks. They can only see what comes out: the probability it says out loud, and the decision it makes.
To test the AI, they used two main "litmus tests," which can be understood through these analogies:
1. The "No Extra Secrets" Test (Conditional Independence)
Imagine the AI is a detective who writes a report.
- The Belief: The report says, "I think the suspect is guilty with 80% certainty."
- The Action: The detective arrests the suspect.
If the report is a complete summary of everything the detective knows, then once you read the "80% certainty" part, you shouldn't learn anything new by looking at the arrest decision. The arrest decision should be fully explained by that 80%.
The Finding: The researchers found that for many AI models, the "80%" number wasn't the whole story. Even after knowing the AI said "80%," looking at what the AI did (arrested or didn't arrest) still gave away extra information about the true state of the patient.
- Analogy: It's like a weather forecaster saying, "There's a 50% chance of rain," but then they are seen carrying a heavy umbrella. If you only heard the "50%," you wouldn't expect the umbrella. The fact that they carried the umbrella means they "knew" something more than they said. The AI's actions revealed hidden information that its spoken words didn't capture.
2. The "Moving Target" Test (Monotonicity)
Imagine you are watching a traffic light.
- If the light is Red (high risk), you stop.
- If the light is Yellow (medium risk), you slow down.
- If the light is Green (low risk), you go.
If the AI is consistent, as the "risk" number it says goes up, its actions should shift predictably toward more cautious choices. If the AI says the risk is 10%, it should act one way. If it says 90%, it should act differently.
The Finding: The researchers checked if the AI's actions moved in the right direction as the numbers changed.
- The Good News: For the strongest, most advanced AI models, the actions did move in the right direction. When the AI said the risk was higher, it acted more cautiously.
- The Bad News: The connection wasn't perfect. Sometimes the AI would say a high risk but act like it was low risk, or vice versa. It was like a traffic light that sometimes flickers between colors even when the car is already moving.
The Verdict: "Close, But Not Perfect"
The paper concludes that while the best AI models are getting better at aligning their words with their actions, they aren't perfectly consistent yet.
- The "Near-Rational" Agent: The researchers proved mathematically that if an AI passes these tests, it acts as if it truly believes what it says. It behaves like a rational person who has a clear internal belief system.
- The Reality Check: Most models failed the "No Extra Secrets" test. This means the AI's internal "brain" holds more information about the patient than it is willing to put into a simple percentage number. The spoken belief is an imperfect summary of what the AI actually knows.
Why This Matters (According to the Paper)
The authors emphasize that in high-stakes fields like medicine, we can't just trust an AI because it says, "I'm 90% sure."
- If the AI's actions don't match its words, we can't trust its explanation.
- The paper provides a way to validate these beliefs. Instead of taking the AI's word for it, we can check if its behavior proves that it actually holds that belief.
In short: The paper teaches us how to catch AI agents when they say one thing but do another. It shows that while the smartest AI models are getting closer to being honest and consistent, they still have "hidden thoughts" in their actions that their spoken words don't fully reveal.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.