← Latest papers
🤖 AI

Superficial Beliefs in LLM Decision-Making

This paper demonstrates that large language models exhibit "superficial beliefs" in decision-making, where their choices are systematically driven by underlying attribute priorities that they can only partially and imperfectly articulate in their self-reported reasons.

Original authors: Gabriel Freedman, Francesca Toni

Published 2026-06-10
📖 5 min read🧠 Deep dive

Original authors: Gabriel Freedman, Francesca Toni

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to figure out how a very smart, but slightly mysterious, robot makes decisions. You ask it to choose between two options, like "Drug A" or "Drug B." The robot picks one and then tells you why it picked it.

The big question this paper asks is: Is the robot actually thinking about the reasons it gives, or is it just making up a story after the fact?

Here is a simple breakdown of what the researchers did and what they found, using some everyday analogies.

The Experiment: The "Taste Test"

The researchers set up a game where the robot had to choose between two profiles (like two different job candidates or two medicines). Each profile had four "stats" (like Safety, Cost, Speed, and Quality), rated as Low, Medium, or High.

They asked the robot to do two things for every choice:

  1. Pick a winner.
  2. Explain its choice by naming the single most important stat that led to the decision.

To see if the robot was telling the truth, the researchers used a "secret decoder ring" (a mathematical model). They watched the robot make hundreds of choices and calculated a hidden "preference score" for each stat based on what the robot actually picked. This is like watching a person eat a thousand meals and guessing their favorite food based on what they order, rather than asking them what they like.

The Discovery: The "Split Personality"

The results were a mix of good news and confusing news.

1. The Robot Does Have a Pattern (The "Hidden Logic")
The "secret decoder ring" worked surprisingly well. When the researchers used the robot's past choices to predict its future choices, they were right about 80% of the time.

  • Analogy: It's like watching a friend order coffee for a month. Even if they don't say it out loud, you can predict they will order a latte tomorrow because they always do. The robot isn't just guessing randomly; it has a hidden, consistent logic.

2. The Robot's "Story" Doesn't Match the "Logic" (The "Superficial Belief")
Here is the twist: When the researchers compared the robot's actual hidden logic with the reasons the robot gave out loud, they didn't match perfectly.

  • The robot's hidden logic said, "I chose Drug A because of Safety."
  • But when asked, the robot often said, "I chose Drug A because of Efficacy."
  • Analogy: Imagine you are driving a car. Your hands are steering left because there is a pothole (the hidden logic). But if someone asks, "Why are you turning left?" you might say, "Because I want to see the view" (the story). You aren't lying maliciously; you just don't have full access to the real reason your hands moved.

The researchers call this "Superficial Belief." The robot behaves as if it has deep beliefs and priorities, but it only has a "superficial" (surface-level) ability to explain them. It can make the right choice, but it can't always articulate the real driver behind that choice.

The "Scorecard" Test

The researchers tried a second way to get the truth. Instead of asking the robot to pick a winner, they asked it to give a "score" (0 to 1) for how important each stat was.

  • Result: This method was more consistent (the robot gave similar scores every time), but it still didn't perfectly match the hidden logic derived from the choices. It was like asking the robot to rate its hunger; it gave consistent numbers, but those numbers still didn't perfectly explain which sandwich it actually wanted to eat.

The "Control" Check

To make sure the robot wasn't just picking words at random, they added "fake" stats that didn't matter at all (like "Packaging Symmetry" for a medicine).

  • Result: The robot almost never picked these fake stats as the reason. This proves the robot is paying attention to the real data; it just isn't always good at explaining which real data point mattered most.

The Bottom Line

The paper concludes that Large Language Models (LLMs) are not just random noise machines, nor are they fully transparent thinkers.

  • They are not random: They follow a hidden, predictable pattern when making decisions.
  • They are not fully transparent: Their spoken explanations are only a partial, imperfect reflection of that hidden pattern.

The Metaphor: Think of the LLM as a very talented chef who can cook a perfect meal every time (the decision). But if you ask the chef, "What was the secret ingredient?" they might guess "Salt" when the real secret was actually "A specific type of pepper." The chef knows how to make the dish work, but they don't have a clear, verbal map of exactly how their own brain is doing it. They have a "superficial belief" in their own cooking process.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →