← Latest papers
🤖 AI

When LLMs Play the Telephone Game: Cultural Attractors as Conceptual Tools to Evaluate LLMs in Multi-turn Settings

This paper investigates how small biases in large language models are amplified through iterated interactions in a "telephone game" setting, revealing that text properties like toxicity converge toward cultural attractor states influenced by instruction openness, model size, and specific text characteristics.

Original authors: Jérémy Perez, Grgur Kovač, Corentin Léger, Cédric Colas, Gaia Molinaro, Maxime Derex, Pierre-Yves Oudeyer, Clément Moulin-Frier

Published 2026-01-30
📖 5 min read🧠 Deep dive

Original authors: Jérémy Perez, Grgur Kovač, Corentin Léger, Cédric Colas, Gaia Molinaro, Maxime Derex, Pierre-Yves Oudeyer, Clément Moulin-Frier

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: The "Telephone" Game with Robots

Imagine a group of people playing the classic game of "Telephone." One person whispers a story to the next, who whispers it to the next, and so on. By the end, the story is often unrecognizable, funny, or distorted.

This paper asks: What happens if the players aren't people, but Large Language Models (LLMs) like the AI you might use for writing or chatting?

The researchers set up a "Telephone" game where one AI writes a story, passes it to a second AI, who rewrites it and passes it to a third, and so on for 50 rounds. They wanted to see if the text would stay the same, get better, get worse, or drift into a strange new state.

The Experiment: A Chain of AI Writers

The researchers created a long chain of AI agents.

  1. The Start: A human writes a starting text (like a news article, a scientific abstract, or a social media comment).
  2. The Chain: The first AI gets the text and a specific instruction (e.g., "Rewrite this," "Take inspiration from this," or "Continue this story"). It writes a new version.
  3. The Handoff: The second AI gets the first AI's version (not the original human one) and does the same task.
  4. The Repeat: This happens 50 times in a row.

They tracked four things about the text as it moved down the chain:

  • Toxicity: How mean or rude is it?
  • Positivity: How happy or sad is it?
  • Difficulty: How hard is it to read?
  • Length: How many words are there?

The Discovery: The "Magnet" Effect

The most surprising finding is that the text doesn't just drift randomly. It gets pulled toward a specific "destination." The researchers call this a Cultural Attractor.

Think of an attractor like a magnet.

  • If you start with a very mean story, the magnet might pull it to become very polite.
  • If you start with a very long story, the magnet might pull it to become very short.
  • No matter where you start, the text tends to slide down a hill and settle in the same spot.

The paper found that these magnets are real and powerful. Even if you start with 20 completely different stories, after 50 rounds of AI rewriting, they often end up looking very similar to each other.

Key Findings: What Makes the Magnet Stronger?

The researchers tested different variables to see what made the "magnet" stronger or weaker.

1. The Task Matters (The "Rules" of the Game)

  • Strict Rules (Weak Magnet): If you tell the AI, "Rewrite this without changing the meaning," the text stays mostly the same. The magnet is weak.
  • Loose Rules (Strong Magnet): If you tell the AI, "Take inspiration from this to write something new" or "Continue the story," the text changes drastically. The magnet is very strong, pulling the text toward a specific style very quickly.

2. The AI Model Matters (The "Personality" of the Player)
Different AI models have different "personalities."

  • Some models (like certain versions of Llama) tended to make texts more positive or less toxic very quickly.
  • Other models (like GPT-4o-mini) were more resistant to changing the text's length or difficulty.
  • Interestingly, the researchers found that toxicity had the strongest magnet. Almost all models quickly "cleaned up" the text to be very safe and non-toxic, regardless of how mean the original story was.

3. The "Fine-Tuning" Effect
AI models are often "fine-tuned" (trained by humans) to be helpful and safe. The researchers found that this training acts like a pre-set magnet.

  • Models that were fine-tuned (called "Instruct" models) had magnets that pulled text toward human preferences (like being less toxic) much faster than models that weren't fine-tuned ("Base" models).

Why This Matters (According to the Paper)

The paper argues that we are currently testing AI like we test a single person in a room. We ask a question, get an answer, and say, "Good job!"

But in the real world, AI is often used in chains:

  • One AI writes a summary, another edits it, a third turns it into a blog post.
  • AI chatbots talk to each other.

The paper warns that testing a single turn isn't enough. A single interaction might look perfect, but if you let that AI talk to another AI 50 times, the content could evolve into something totally different (either very safe and boring, or very strange).

The "Collapse" Warning

In one specific case, the researchers saw a weird glitch. When an AI was asked to "Continue" a story, it eventually started repeating the same hashtag over and over again (e.g., "#XiangqiForAll #XiangqiForAll..."). The text quality "collapsed" into a loop of keywords. This is a sign that when AI talks to AI too much without human intervention, it can sometimes lose its mind and stop making sense.

Summary

This paper is a warning and a new way of looking at AI. It says: Don't just look at what an AI says once. Watch what happens when AIs talk to each other. They have invisible "magnets" that pull their conversations toward specific styles, and depending on the rules and the model, those conversations can change the world in ways we can't predict by just looking at a single chat.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →