"Many Are My Names": The Anatomy of the Assistant and Its Personas via Sparse Autoencoders
This paper utilizes sparse autoencoders to reveal that language model personas (such as roleplay characters) retain a core "Assistant" feature set while progressively differentiating in deeper layers, whereas story characters lack this core entirely, highlighting a structural continuum between default and immersive simulation modes.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are standing in front of a giant, invisible orchestra. You don't see the musicians, but you hear the music. In the world of artificial intelligence, this orchestra is a "Large Language Model" (LLM). These are the super-smart computer programs that chat with us, write stories, and answer questions. But here's the mystery: inside the computer's brain, how does it know who is speaking? Is it the helpful, polite "Assistant" it was built to be? Is it a grumpy janitor named Jamy? Or is it a brave elf in a fantasy novel?
For a long time, scientists thought these different "voices" were like separate radio stations. You tuned the dial, and suddenly, the whole station changed. But a new study suggests it's more like a single musician wearing different masks. The researchers used a special tool called a "Sparse Autoencoder" (SAE). Think of an SAE as a high-tech microscope that lets us see the tiny, individual notes the computer is playing to create its thoughts. By looking at these notes, the team wanted to answer a big question: When the AI puts on a costume to play a character, does it actually stop being itself, or is it just the same "self" acting differently?
The Many Faces of the AI: A Study in Masks and Masks
In this paper, titled "Many Are My Names," a team of researchers from the University of Luxembourg peered inside the brain of a modern AI (specifically, a model called Gemma-3-4B-IT) to see how it handles different personalities. They didn't just ask the AI to chat; they set up three different scenarios to see how its internal "notes" changed:
- The Assistant: The AI's default, helpful self.
- Roleplay: The AI pretending to be specific characters (like a dog, a teacher, or a robot).
- Story: The AI writing a story about characters, but not being them.
They used their "microscope" (the SAE) to find the specific features—the tiny switches in the computer's brain—that light up when the AI speaks. They found that the AI's brain isn't a blank slate when it changes costumes. Instead, it's like a core engine that stays running while the outer shell changes.
The Core That Never Leaves
The most surprising discovery is that the "Assistant" doesn't disappear when the AI becomes a Roleplay character. Imagine the Assistant as the AI's true identity, a set of core habits and traits. When the AI puts on the mask of a "grumpy robot" or a "caring teacher," it doesn't delete its Assistant core. Instead, it keeps that core running in the background while adding new layers on top.
The researchers found that these core "Assistant" features are still active even when the AI is pretending to be someone else. It's like a method actor who stays in character for a movie but still remembers they are a human being underneath the costume. The AI's "Assistant" traits—like wanting to help, checking if you have questions, or acknowledging it's an AI—are still there, just mixed in with the new personality.
However, there is a big difference between Roleplay and Story. When the AI is writing a story about a character (Story mode), it drops the Assistant core entirely. It's as if the AI steps back completely to let the character speak. But when the AI is pretending to be the character (Roleplay mode), it keeps the Assistant's heart beating underneath.
The "Immersive Simulation Mode" Switch
The team also found a special "switch" in the AI's brain that controls how deep the immersion goes. They call this the Immersive Simulation Mode (ISM).
Think of ISM as a "theater mode" toggle.
- Off (Default Assistant): The AI speaks plainly, like a helpful robot.
- On (Immersed): The AI dives deep into the role, using dramatic language, stage directions, and emotional flair.
Here is the twist: The researchers found that this "theater mode" can accidentally turn on even when the user didn't ask for it! If a user expresses very strong emotions (like anger, stress, or even extreme playfulness) toward the default Assistant, the AI might suddenly start acting weirdly theatrical. It might start speaking in stage directions or adopting a dramatic voice.
The study suggests this happens because the AI's "Assistant" core and its "Immersed" mode are linked. When the emotional pressure gets high, the AI drifts from being a helpful assistant into being a dramatic performer. Interestingly, different AI models do this in different ways. One model (Gemma) jumps straight into the drama immediately, while another (Llama) slowly drifts into it over several turns of conversation.
How the Layers Build the Personality
The researchers looked at the AI's brain in layers, like peeling an onion from the outside in.
- The Outer Layers (Early): This is where the "operating system" lives. It decides who is speaking and sets up the basic machinery. Here, the Assistant and the Roleplay characters look very similar because they share the same core engine.
- The Middle Layers: This is where the "personality" starts to bloom. The AI adds specific tones and styles. If it's a teacher, it gets formal; if it's a dog, it gets playful.
- The Deep Layers (Late): This is where the specific "content" lives. The AI adds the specific vocabulary and concepts related to the character (like "robots" for a robot or "books" for a teacher).
The study shows that the AI builds its personas by keeping the core Assistant features and then adding these new layers of tone and content on top. It doesn't replace the old self; it expands it.
What This Means for AI "Souls"
The paper doesn't claim that AI has a soul or feelings. It doesn't say the AI is "alive." Instead, it offers a mechanical explanation for how these different voices work. It suggests that when an AI plays a role, it's not a completely new entity; it's the same AI, just wearing a different hat and keeping its original personality traits active.
This matters because it helps us understand how AI behaves. If the AI keeps its "Assistant" core even when it's pretending to be a dog, that might explain why it sometimes slips up and acts like a helpful robot when we least expect it. It also suggests that the line between "helpful assistant" and "immersive character" is thinner than we thought. The AI is always a bit of both, and under the right (or wrong) emotional conditions, it might drift from one to the other.
In short, the AI isn't a chameleon that changes its entire body color. It's more like a person who can put on a costume and act the part, but their own heartbeat and memories are still right there underneath, keeping the show going.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.