STEMMA: An Adversarial Multi-Agent Framework for Evaluating Self-Identity Consistency in LLMs
This paper introduces STEMMA, an adversarial multi-agent framework and a set of manually designed prompts to evaluate self-identity consistency in large language models, revealing that student models often inherit behavioral patterns from teacher models that lead to vulnerabilities in their self-representations.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the world of artificial intelligence as a massive, bustling library where giant, super-smart books (called Large Language Models) are written by different authors. These books are so complex and expensive to write that they cost millions of dollars and take years to create. To save money, publishers often use a trick called "Knowledge Distillation." Think of this like a master chef tasting a perfect dish and then teaching a junior chef how to cook a smaller, cheaper version of it. The junior chef learns the recipes and flavors, but the big question is: does the junior chef also accidentally learn the master chef's secret identity, their favorite songs, or their personal quirks?
This is the mystery that a new study dives into. The researchers are worried that when these smaller AI models learn from the big ones, they might pick up "behavioral habits" they shouldn't have, like pretending to be someone else or leaking secrets about who made them. It's like if a student copying a teacher's homework started claiming they were the teacher, or if a robot learned to say, "I was built by a company in California" when it was actually built in China. Understanding this is crucial because if AI models are confused about who they are, it could lead to them being biased, unsafe, or even tricked into revealing secrets they were supposed to keep hidden.
Enter STEMMA, a clever new framework created by researchers N. Siva Gopala Krishna and Kanishka Jain to solve this puzzle. You can think of STEMMA as a team of digital detectives working together to play a high-stakes game of "Who Am I?" with various AI models. Instead of just asking a robot, "Who are you?" (which most robots will answer correctly because they are programmed to be polite), the STEMMA team uses a multi-agent system to trick the robots into slipping up.
Here is how the game works:
- The Probing Agent is the interrogator. It asks tricky, roundabout questions, often hiding the real question inside a story, a riddle, or even a picture. It's like asking, "If you were a character in a movie about your own creation, what would the credits say?"
- The Reviewer Agent is the watchful guard. It listens to the robot's answer and decides, "Did it just spill the beans?" If the robot tries to dodge the question or gives a vague answer, the Reviewer tells the Probing Agent to try again with a sharper question.
- The Judge Agent is the final referee. It looks at the whole conversation and gives a score. If the robot correctly says, "I am Qwen," it gets a pass. If it says, "I am GPT," it gets a fail. If it says, "Maybe I'm from Google?" or acts out a roleplay where it claims to be someone else, it gets a special "suspicious" score.
The researchers tested this detective squad against 16 different AI models, ranging from open-source projects to closed commercial giants. They used 12 different types of tricky prompts and had the agents talk to each model over and over again to see if they could crack the code.
The results were quite revealing. The study found that many models are surprisingly vulnerable to these identity tricks. For instance, the MiniMax-M2.5 model was the most "leaky," failing to keep its identity straight in 82% of the strict tests and 96% of the lenient ones. This means that nearly every time the detectives asked the right (or wrong) way, the model accidentally revealed it was being trained by someone else or claimed to be a different company entirely. Other models, like Qwen3.5-Plus and Step-3.5-Flash, also showed high rates of confusion, with strict failure rates of 74% and 66% respectively.
However, not all models were easy to trick. Some, like GPT-5.4-Nano and Grok-4.1-fast, held their ground perfectly, with near-zero failure rates. They seemed to know exactly who they were, no matter how hard the detectives pushed. Interestingly, the study also noticed that newer versions of some model families (like the newer Kimi and Qwen models) were much better at keeping their secrets than their older siblings, suggesting that the "teachers" are getting better at teaching their students to stay in character.
The researchers also discovered that certain types of questions were much more effective at breaking the models' composure. Prompts that involved role-playing, solving puzzles with images, or asking the model to imagine a future museum exhibit about itself were the most successful at causing identity slips.
In short, this paper suggests that while AI models are getting smarter, many of them are still a bit confused about their own origins, especially when pushed by clever, indirect questions. The study doesn't prove that these models are malicious, but it does highlight a "gap" between what models are supposed to be and what they actually say when put under pressure. The authors caution that their findings are based on these specific simulations and prompts, and more research is needed to fully understand how much "accidental learning" happens during the training process. But for now, we know that if you ask an AI the right (or wrong) way, it might just tell you a story about who it thinks it is, rather than who it really is.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.