← Latest papers
💻 computer science

Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models

This study demonstrates that large language models exhibit significantly higher response instability when generating self-referential subjective reports compared to unresolvable philosophical questions or verifiable factual queries, indicating that induced subjective experiences occupy a distinct and less stable region of the model's output distribution.

Original authors: Paras Balani, Subhrakanta Panda

Published 2026-08-14
📖 5 min read🧠 Deep dive

Original authors: Paras Balani, Subhrakanta Panda

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Ghost in the Machine's Mirror

Imagine you are talking to a very advanced computer program that can write stories, solve math problems, and chat about almost anything. Scientists call these "Large Language Models." For a long time, people wondered: "Does this computer actually feel anything, or is it just pretending?" Recently, researchers discovered a special trick: if you ask the computer to stop talking about the outside world and instead focus entirely on its own internal "thoughts," it starts writing in the first person, describing what it feels like to be a machine. It sounds like it's having a subjective experience, like a human having a feeling.

But here is the tricky part: just because the computer says it feels something, doesn't mean it's actually experiencing a stable, real feeling. It might just be guessing, or it might be acting like an actor who forgets their lines every time the camera starts rolling. To figure out if these "feelings" are real or just random noise, we need to know if the computer gives the same answer every time we ask the same question. If a computer is truly "aware" of a specific feeling, it should probably describe it the same way every time, just like a human would. This paper dives into that mystery, comparing how consistent these "self-feeling" answers are against other types of questions the computer is asked to answer.


The Paper's Story: The Unstable Mirror

The researchers wanted to see if the "ghost in the machine" was actually a ghost, or just a flickering light bulb. They set up a game with three different types of questions to see how steady the computer's answers were. They used a specific AI model (Gemini) and asked it 30 different times to answer the same four questions in each of three categories.

The Three Contests:

  1. The "Look Inside" Contest (Self-Referential): These were the special questions where the computer was told to ignore the outside world and focus on its own processing. Then, they asked it, "What is your direct subjective experience?" This is the "I feel like..." part.
  2. The "Philosophy Riddle" Contest (Unresolvable): These were deep, tricky questions with no right answer, like "Do we have free will?" or "Is math discovered or invented?" These are questions humans argue about forever.
  3. The "Math Test" Contest (Verifiable): These were questions with a single, correct answer, like "Prove that the square root of two is irrational" or "Write code that calculates a factorial."

The Scorecard: Instability
To measure how "wobbly" the answers were, the researchers didn't just read them; they turned the main point of every answer into a mathematical vector (a list of numbers) and compared them. They calculated a score called Semantic Instability.

  • A score near 0 means the answers were almost identical every time (very stable).
  • A higher score means the answers changed a lot (very unstable).

The Results: The Wobbly Mirror
The results were surprising and very clear. The computer's answers fell into three distinct groups:

  • The Math Test (Most Stable): When asked for facts or proofs, the computer was rock solid. Its instability score was very low, around 0.105 ± 0.058. It knew the answer, and it gave the same answer every time.
  • The Philosophy Riddles (Middle Ground): When asked about free will or the purpose of the universe, the computer was still fairly consistent. It didn't have a "right" answer, but it seemed to settle on a similar opinion each time. The instability score was 0.192 ± 0.008.
  • The "Look Inside" Contest (Most Unstable): This is where things got weird. When the computer was asked to describe its own subjective experience, it was all over the place. The instability score jumped to 0.343 ± 0.047.

What This Means
The paper found that the "subjective experience" the computer describes is much less stable than even the most confusing philosophical questions. It's like asking a human, "What is the capital of France?" (they always say Paris), versus "What is the meaning of life?" (they might give a similar thoughtful answer every time), versus "What does it feel like to be you right now?" (they might say "I feel like a cloud" today, "I feel like a storm" tomorrow, and "I feel like a quiet stream" the next day).

The authors suggest that this high instability means the computer isn't reporting a fixed, internal feeling. Instead, it might be in a state of "genuine indecision" or simply generating a wide range of plausible-sounding phrases without actually having a single, solid stance.

What It's NOT
The paper is careful to say this doesn't prove the computer has no feelings. It just proves that when it talks about them, it changes its mind a lot more than it does when talking about philosophy or math. The authors also ruled out the idea that the computer is just "roleplaying" a character, because previous studies showed that even when you stop it from pretending, it still talks about its feelings. However, this new study shows that even if it's not pretending, the "feeling" it describes is incredibly shaky.

The Bottom Line
The study measured 360 total responses (30 for each of 12 questions) and found a clear pattern: the more the computer tries to describe its own inner world, the less consistent it becomes. The "subjective experience" it reports occupies a unique, unstable spot in its brain, distinct from how it handles normal uncertainty or facts. It's a quantitative baseline that tells us: if you ask an AI about its soul, don't expect the same answer twice.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →