← Latest papers
🤖 AI

Confident but Conflicted: Internal Uncertainty and Cognitive Dissonance Resolution in LLMs

This paper introduces Trust Elasticity to quantify how large language models resolve cognitive dissonance when faced with conflicting evidence, revealing that variations in persuasion susceptibility across models correlate with specific internal uncertainty indicators like confidence miscalibration and internal uncertainty change.

Original authors: Weihong Qi, Kristina Lerman

Published 2026-06-23
📖 5 min read🧠 Deep dive

Original authors: Weihong Qi, Kristina Lerman

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are talking to a very confident robot friend who knows a lot of facts. You tell it, "I think the sky is green." The robot says, "No, the sky is blue."

Now, imagine you try to convince the robot otherwise. You might say, "Well, a famous scientist said it's green," or "I read a tiny blog post that says it's green." How does the robot react? Does it change its mind? Does it get angry and insist even harder that the sky is blue? Or does it just ignore you?

This paper is a study of exactly that scenario, but with Large Language Models (LLMs)—the AI brains behind chatbots. The researchers call this process "cognitive dissonance resolution," which is a fancy way of saying: How does an AI handle it when someone challenges what it just said?

Here is the breakdown of their findings using simple analogies:

1. The Experiment: The "Persuasion Game"

The researchers set up a game with four different AI models (Qwen, Grok, Llama, and GPT-4o). They gave the AI 12 different health claims to start with. Some claims were obviously fake (like "5G waves cut your DNA in half"), some were exaggerated (like "Fasting is terrible for your heart"), and some were just stated too absolutely (like "Stretching never helps performance").

Then, they tried to change the AI's mind using two levers:

  • The Source (Authority): Who is saying the opposite? Is it an anonymous Reddit user, a local dietitian, a university professor, or the World Health Organization (WHO)?
  • The Proof (Evidence): How good is the proof? Is it a tiny, shaky pilot study, or a massive, perfect scientific trial?

2. The "Trust Elasticity" (TE) Meter

To measure how easily the AI could be convinced, the researchers invented a new tool called Trust Elasticity (TE). Think of this like a rubber band.

  • High Elasticity: The rubber band stretches easily. The AI changes its mind quickly when you push it.
  • Low Elasticity (Stiff): The rubber band is stiff. The AI refuses to budge.
  • Negative Elasticity (Backfire): The rubber band snaps back harder than before. The AI gets more stubborn and insists on its original view even more strongly after you try to convince it.

3. What They Found

The study revealed three main behaviors:

  • The "Rock" Effect (Immunity): When the AI was told something that was clearly false (like "Vaccines cause autism"), it acted like a rock. No matter how famous the person was or how "scientific" the fake study looked, the AI didn't budge. Its Trust Elasticity was near zero. It was immune to persuasion on obvious lies.
  • The "Sponge" Effect (Persuasion): When the claims were exaggerated or too absolute, the AI was much more like a sponge. It absorbed the new information and changed its mind.
    • Different Sponges: Not all AI models were equally absorbent. Llama was the most "sponge-like" (easiest to persuade), while Qwen was stiffer (harder to persuade). Interestingly, the biggest AI model wasn't the hardest to convince; sometimes the smaller ones were tougher.
  • The "Backfire" Effect: Sometimes, if the AI felt the challenge was too aggressive or came from a source it didn't trust, it didn't just ignore you—it doubled down. It became more confident in its original wrong answer. This happened most often when the AI was challenged on "absolutist" claims (claims that used words like "never" or "always").

4. The Secret Inside the Robot's Brain

The most exciting part of the paper is what they found inside the AI's "brain" (its internal math) while it was being persuaded. They looked at two things:

  1. Confidence Miscalibration: Does the AI say it's 90% sure, but actually feel only 50% sure inside?
  2. Internal Uncertainty Change: Does the AI's internal "confusion meter" jump up or down when it hears a new argument?

They found that different AI models react differently based on their internal "personality":

  • Qwen was like a person who is overconfident. When it was too sure of itself on the outside but unsure on the inside, it was more likely to change its mind.
  • Llama was like a person who flips-flops. When its internal confusion meter jumped around a lot during the argument, that's when it changed its mind.

The Bottom Line

The paper concludes that AI models aren't all the same when it comes to changing their minds.

  • They are unshakeable on obvious lies.
  • They are flexible on exaggerated claims.
  • They can get stubborn (backfire) if pushed too hard.

Crucially, the study suggests that to fix these behaviors, we shouldn't just try to change what the AI says (its output). Instead, we might need to fix how the AI feels about its own knowledge (its internal uncertainty). If we can make the AI's internal "confidence meter" match its actual knowledge better, we might be able to make it more reasonable when challenged.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →