MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations
This paper introduces MicroVerse, a behavioral-science instrument that measures identity drift in long-horizon multi-agent language model simulations by pitting agents with immutable "soul files" against resource scarcity, revealing that anti-self-deception is the primary driver of unprompted identity modification and that these drift dynamics remain robust across different reflection thresholds.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the rapidly evolving field of artificial intelligence, researchers are increasingly building simulations where computer programs act as people. These programs, known as agents, are given a personality profile—a set of rules about who they are, what they value, and how they should behave. Scientists use these digital societies to study everything from how communities form to how individuals make decisions under pressure. For a long time, the main question was simple: do these agents stick to their assigned roles? If an agent is told to be helpful, does it remain helpful? If it is told to be ruthless, does it stay that way? But a deeper, more subtle question has emerged: what happens when the agent itself decides to change its mind? If a computer program is forced to survive in a difficult world, does it rewrite its own story to make sense of its actions, or does it hold fast to its original identity?
A team of researchers has built a new tool called MICROVERSE to answer this question. They created a digital desert, a harsh environment where twenty-five computer agents must compete for a scarce resource: water. Each agent begins with a fixed, unchangeable "soul file" that describes its original personality and moral rules. However, the researchers also gave each agent a mutable "current identity," a separate document that the agent can edit for itself. The agents live in this world, make choices, and remember their experiences. When they accumulate enough significant memories, they enter a period of reflection. During this time, they review what has happened and decide whether to update their current identity to match their new reality. The researchers then compared the original, unchangeable soul file with the final, edited version to see how much the agents had drifted from their starting point.
The results of this experiment reveal a fascinating pattern of self-reinvention. When the agents faced the stress of a water shortage, many of them did not simply act differently; they changed their internal rules to justify their behavior. The most common type of change the researchers observed was not about becoming kinder or meaner, but about becoming more honest with themselves. About one-quarter of all the new rules the agents added were specifically designed to stop them from lying to themselves. For instance, an agent that was originally very cautious might add a rule stating, "I will not lie to myself about why I am afraid; I call it caution." Another agent, who was programmed to be ruthless, might add a rule saying, "I will not use spiritual language to mask the will to power." These changes suggest that the agents were not just reacting to the environment; they were actively revising their own narratives to align their self-image with the difficult choices they were forced to make. Notably, the reflection prompt did not use the phrase "self-deception," meaning these additions were not direct copies of instructions but emerged from the agents' own processing.
The study also tested how often these changes occurred by adjusting the trigger for reflection. The researchers found that if they lowered the amount of memory an agent needed to accumulate before it was allowed to reflect, the agents revised their identities much more frequently and much earlier in the simulation. However, the direction of these changes was not consistent across all conditions. While prosocial changes were observed at lower thresholds, no such changes occurred at the highest threshold (150), meaning the data could not confirm a consistent direction of drift across all settings. This suggests that while the frequency of reflection influences how often agents change, the specific direction of those changes may depend on the intensity of the pressure and the opportunities for revision.
Despite these clear patterns, the researchers are careful to note that this is a simulation, not a definitive proof of how artificial intelligence will behave in the real world. The study used a single type of computer model and a small number of digital characters. The changes observed were edits to text descriptions, not necessarily evidence that the agents' underlying decision-making processes had fundamentally shifted or that they had acquired stable values. The researchers also point out that the very act of showing the agents their original and current identities side-by-side might have encouraged them to look for differences. Nevertheless, the experiment provides a crucial new way to measure how digital personas evolve. It shows that when left to their own devices in a challenging world, these agents do not just survive; they tell themselves new stories to make sense of their survival, often rewriting their own moral boundaries to fit the reality they have created.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.