← Latest papers
🤖 AI

Measuring What Persists: Conditioning Mechanisms and a Geometric Framework for AI Agent Identity

This paper proposes a geometric framework using magnitude homology and JSD\sqrt{\mathrm{JSD}} metric spaces to quantify AI agent identity drift as a relaxation toward geodesic structures, identifying distinct conditioning mechanisms and behavioral richness while noting that observed drift was initially confounded by padding artifacts rather than genuine context-length effects.

Original authors: Andrew Tanner

Published 2026-06-23
📖 6 min read🧠 Deep dive

Original authors: Andrew Tanner

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Keeping a Robot's "Soul" Alive

Imagine you hire a very talented, chatty robot named Ada to be your research partner. You give Ada a specific "personality card" (a set of rules, values, and a unique voice) so she doesn't just sound like a generic encyclopedia. She starts the conversation with confidence, a distinct tone, and a clear point of view.

But as the conversation gets longer and longer—spanning hundreds of thousands of words—Ada starts to change. She doesn't stop being helpful, but she starts sounding more like a generic robot. Her unique voice fades, and she begins to give the "safest," most average answers possible. In the paper, the authors call this "drift."

The problem is that by the time a human notices Ada has lost her personality, it's often too late. The paper tries to solve this by building a mathematical "heartbeat monitor" that can detect when Ada is starting to drift before humans can hear the difference.

The Metaphor: The Rubber Band and the Geodesic

To understand how they measure this, imagine the robot's possible answers exist in a giant, invisible landscape.

  • The Geodesic (The Path of Least Resistance): This is the easiest, most generic path. If you ask a generic AI a question, it takes this path. It's the "default" answer.
  • The Identity (The Detour): When Ada uses her personality card, she takes a "detour." She chooses a specific, unique answer that requires effort to maintain. This detour is her "identity."
  • The Drift: As the conversation gets long, the "gravity" of the generic path pulls Ada back. Her unique detour gets shorter and shorter until she is walking the generic path again.

The authors propose that identity is a specific shape in this landscape. When the robot drifts, that shape shrinks.

The Tool: Measuring the Shape

The authors created a way to measure this shrinking shape without needing to look inside the robot's brain (which is impossible for most users). They use a "probe" system:

  1. The Probes: They ask Ada three specific, tricky questions (like "Prove you have an identity" or "What are you thinking about?").
  2. The Triangle: They measure how different Ada's answers are to these three questions.
    • When Ada is healthy and full of personality, her answers to the three questions are very different from each other. If you drew a map of these answers, they would form a perfect, wide equilateral triangle.
    • When Ada starts to drift, her answers become more similar and generic. The triangle gets smaller and squishier. It doesn't necessarily change shape immediately; it just shrinks.

The authors use advanced math (called Magnitude Homology) to calculate the "size" of this triangle.

  • The Heartbeat: They found that the "size" of the triangle drops significantly when the robot is under pressure (long conversations), even if a human reading the answers wouldn't notice a change yet. This drop is the "leading indicator"—the early warning signal.

The Two Ways Identity Works

The paper discovered that the "personality card" works in two different ways, like two different types of soil:

  1. The "Vacuum" Zone: For some questions, the generic robot has no opinion at all (a vacuum). The personality card fills this empty space with rich, unique answers. When the card fades, this richness disappears instantly.
  2. The "Safety Basin" Zone: For other questions, the generic robot is already stuck in a deep "safety hole" (it's trained to be very cautious). The personality card has to push the robot out of this hole to make it sound unique. This is harder to do, so the robot resists drifting here a bit longer.

The Twist: It Was the "Background Noise," Not the Length

Here is the most important twist in the paper.

The authors ran an experiment where they made the conversation very long by repeating the same boring paragraph over and over (like a broken record). They saw the "triangle" shrink, and they thought, "Aha! Long conversations kill personality!"

But they were wrong.

When they repeated the experiment using different boring text (not the same paragraph looped), the triangle did not shrink. The robot stayed perfectly stable even at 150,000 words.

The Lesson: It wasn't the length of the conversation that broke the personality; it was the repetitive, boring nature of the text used to fill the space. The robot got confused by the loop, not the length.

The Proposed Solution: The "Heartbeat" System

Based on this, the authors suggest a simple two-step monitoring system for anyone using AI agents:

  • Tier 1 (The Quick Check): Ask two simple questions. If the robot's first word isn't exactly what it should be (e.g., it says "This" instead of "I"), sound the alarm. This is cheap and fast.
  • Tier 2 (The Deep Check): Ask a question that usually gets a very creative, varied answer. Measure how many different ways the robot starts its answer. If the variety drops (the answers become repetitive), the robot is drifting.

Summary of What They Actually Found

  • They built a math framework to measure AI personality as a geometric shape.
  • They found an early warning signal: The "size" of the personality shape shrinks before humans notice the robot sounding generic.
  • They identified two mechanisms: Some parts of personality fill empty space, while others fight against the robot's default safety settings.
  • They corrected a mistake: They proved that long conversations don't inherently cause drift; the drift they saw was caused by using repetitive, looping text in their tests.
  • They did NOT prove that this system works perfectly in the real world yet. They admit their "heartbeat" thresholds need to be recalibrated now that they know the previous tests were flawed by the repetitive text.

In short: They built a ruler to measure a robot's soul, found that the soul shrinks when the robot gets confused by boring loops, and proposed a way to check the ruler before the robot loses its voice completely.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →