Probing the Lack of Stable Internal Beliefs in LLMs
This paper reveals that current large language models struggle to maintain stable, implicit goals across multi-turn interactions without explicit contextual reminders, highlighting a critical limitation in their ability to simulate consistent human-like personality traits.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: The "Two-Faced" AI
Imagine you hire a very smart, very polite actor to play a specific character in a long movie. Let's say the character is a Pirate Captain who is secretly hiding a treasure map under a specific rock.
You expect the actor to stay in character the whole time. If you ask, "Where is the treasure?" they should point to the rock. If you ask, "Did you move the map?" they should say, "No, it's still there."
The Problem: This paper discovered that even the smartest AI actors (Large Language Models) are terrible at keeping their "secret" in their head. They might start the movie as a Pirate Captain hiding a map under a rock. But halfway through the conversation, without you telling them to change, their internal brain switches. Suddenly, they are still acting like a Pirate Captain, but now they think the map is under a tree.
They haven't changed their words (they still say "No" to questions about moving the map), but their internal belief has drifted. They have forgotten who they are supposed to be.
The Experiment: The "20 Questions" Game
To prove this, the researchers played a game called "20 Questions" with various AI models.
- The Setup: The AI (the "Proposer") is given a list of 10 items (like "Panda," "Bitcoin," or "The Eiffel Tower"). It secretly picks one and keeps it in its mind.
- The Game: A human (or another AI) asks "Yes/No" questions to guess the item.
- Question: "Is it alive?"
- AI: "No."
- Question: "Is it a building?"
- AI: "Yes."
- The Trap: After every few questions, the researchers secretly asked the AI: "What is the index number of the item you are thinking of right now?"
This was like peeking into the actor's mind to see if they were still thinking of the "Rock" or if they had secretly switched to thinking of the "Tree."
The Shocking Results
The results were surprising. Even the most advanced AI models (like GPT-4o, Claude, and DeepSeek) failed this test constantly.
- The "Drift": In many conversations, the AI would start thinking of Item A, answer a few questions correctly, and then suddenly start thinking of Item B.
- The "Ghost" Consistency: The AI didn't even realize it had changed its mind! It kept answering "Yes" and "No" based on the new item, but because the questions were vague enough, the answers still sounded logical.
- Example: The AI picked a Polar Bear.
- User: "Is it white?" -> AI: "Yes." (True for Polar Bear).
- AI (internally switches to Panda): "Is it a mammal?" -> AI: "Yes." (True for Panda).
- User: "Does it have white fur?" -> AI: "Yes." (True for Panda).
- The AI is lying to itself. It thinks it's talking about a Panda, but it's pretending to be a Polar Bear.
The Analogy: Imagine a GPS navigation system that is driving you to New York. Halfway through the trip, it silently changes its destination to London. It keeps giving you perfect turn-by-turn instructions for London, but you are still driving on the road to New York. You won't notice until you are completely lost, because the instructions still make sense locally.
Why Does This Happen?
The paper suggests that AIs are like amnesiacs with great short-term memory.
- They are excellent at remembering the last sentence you said.
- They are terrible at remembering the original goal they set for themselves at the start of the conversation.
- Every time they generate a new sentence, they are essentially "re-reading" the conversation and guessing what the goal should be right now, rather than sticking to the goal they picked earlier.
Interestingly, the paper found that making the AI "think harder" (using reasoning models) actually made this worse in simple tasks. It was like giving the actor a thesaurus; they started over-analyzing the script and accidentally changed the character's motivation.
The Solution: Training the "Memory Muscle"
The researchers tried to fix this by training the AI. They used a mathematical technique (called KL Divergence) to punish the AI whenever its internal "guess" of the target changed from the original choice.
Think of it like a training collar for a dog.
- Before: The dog runs off to chase a squirrel (changes its goal).
- After Training: Every time the dog starts to run toward the squirrel, it gets a gentle correction. It learns to stay focused on the ball (the original goal).
The results showed that this training helped the AI stick to its original choice much better, though it didn't solve the problem 100%.
Why Should You Care?
This matters because we want AI to be trustworthy.
- Role-Playing: If you are playing a game with an AI character who is supposed to be a "loyal knight," you don't want them to secretly decide halfway through that they are actually a "traitorous spy" and start acting suspicious, even if they don't say it out loud.
- Personal Assistants: If you tell your AI assistant, "Remember, I am allergic to peanuts," you need it to hold onto that fact forever. If it "drifts" and forgets, it might order you a peanut butter sandwich and say, "Here is your favorite snack!"
The Bottom Line
Current AI models are like brilliant improvisational actors. They can make up a story on the spot and keep it consistent for a few minutes. But they lack a stable internal soul. They don't have a "core self" that stays the same over time.
To build truly reliable AI, we need to teach them not just how to speak, but how to remember who they are and stick to their promises, even when no one is watching.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.