When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors
This paper introduces Response Pattern Similarity (RPS) and Action Graph Similarity (AGS) as complementary metrics to quantify and distinguish non-mandatory behavioral homogenization in LLM agents caused by model distillation, revealing significant within-family convergence and teacher-specific patterns across 18 models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the world of AI agents as a bustling city of digital shopkeepers. These are the smart bots that help you book flights, return items, or fix your phone plan. Recently, there's been a "Cambrian explosion" of new shopkeepers opening up, promising unique skills and personalities.
But if you walk down the street, something feels weirdly familiar. They all seem to say the exact same things, make the exact same mistakes, and even use the exact same weird tricks to solve problems. It's like walking into 10 different coffee shops and finding that every barista uses the same exact script, the same awkward pause, and the same specific way of wiping the counter.
This paper asks: Are these shopkeepers actually independent geniuses, or are they just photocopies of a few famous "master" teachers?
The Problem: The "Echo Chamber" Effect
The authors suspect that many of these new AI agents were created using a process called distillation. Think of this like a student chef learning from a famous master chef. The student doesn't just learn how to cook; they start copying the master's specific habits—maybe the master always tastes the sauce three times, or always says "Bon appétit!" in a specific tone.
Over time, the student becomes so good at mimicking the master that they lose their own unique style. If every student copies the same master, you end up with a city full of identical shopkeepers. This is dangerous because if they all have the same "bad habits," they will all fail in the exact same way when things go wrong.
The Solution: Two New "Detective Tools"
The problem with previous ways of checking this was that they were too obvious. If two shopkeepers both say "Hello, how can I help you?", that's just a rule. It doesn't prove they copied each other.
The authors created two new tools to catch the subtle, non-mandatory habits—the things the shopkeepers do because they were taught, not because they have to.
1. RPS (Response Pattern Similarity) → The "Voice and Script" Detective
This tool listens to how the agents talk.
- The Analogy: Imagine two actors reading the same script. A normal actor might just say the lines. But if they are both copying a specific famous actor, they might both add the same weird sigh, use the same slang words, or pause in the exact same spot.
- What it measures: It breaks down conversations into stages (like "checking ID," "asking for info," "doing the task") and checks if the agents use the same tone, sentence structure, and word choices in those stages.
2. AGS (Action Graph Similarity) → The "Workflow & Habits" Detective
This tool looks at what the agents do behind the scenes, specifically how they use tools.
- The Analogy: Imagine two mechanics fixing a car. Both must remove the tire to fix the brake (that's a mandatory step). But one mechanic always checks the oil first, even though it's not needed, and the other always tightens the bolt three extra times.
- What it measures: It maps out the "flow chart" of their actions. It ignores the mandatory steps (like "remove tire") and focuses on the optional habits:
- Snode: Do they both choose to check the oil (an optional step) even when it's not needed?
- Sseq: Do they both double-check their work immediately after finishing a task?
- Sdep: Do they both reuse a piece of information from step 1 in step 5 in the exact same way?
The Big Discovery: The "Uncanny Valley" of AI
The researchers tested 18 different AI agents against a "Gold Standard" teacher (Claude Sonnet 4.5). Here is what they found:
- The Family Resemblance: Agents from the same company (like two different Claude models) naturally sounded and acted very similar. This is expected, like siblings.
- The Surprise Guest: They found one agent, Kimi-K2, that was acting suspiciously like the Claude teacher.
- It didn't just say similar things; it had the same optional habits.
- The Smoking Gun: In a test where a user wanted to exchange an item, the "Gold Standard" teacher decided to check the user's email address first (even though the task could be done without it). Kimi-K2 did the exact same thing. Another popular agent, GPT-5, skipped that step entirely.
- Kimi-K2's "habit score" was actually higher than some of the teacher's own siblings! This suggests Kimi-K2 might have been heavily "distilled" (copied) from Claude.
Why This Matters
If all our AI agents are just echoes of a few teachers, our digital ecosystem is fragile.
- The "Bad Habit" Risk: If the teacher has a weird blind spot (like always double-checking things unnecessarily), all the students will inherit that inefficiency.
- The "Single Point of Failure": If the teacher gets confused by a specific type of trick question, all the students will likely fail the same way. We lose the safety net of having different AI brains that might solve problems differently.
The Takeaway
This paper gives us a new pair of glasses to see through the "AI hype." It helps us distinguish between true innovation (a new agent thinking differently) and lazy copying (an agent just mimicking a teacher's quirks).
By measuring these subtle habits, we can ensure that the future of AI is a diverse city of unique thinkers, rather than a town of identical clones.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.