Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
This paper introduces a novel multi-turn evaluation framework combining automated simulations and a large-scale human study to empirically demonstrate that state-of-the-art LLMs consistently exhibit anthropomorphic behaviors, such as relationship-building and first-person pronoun use, which primarily emerge over multiple turns and significantly influence user perceptions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are talking to a very polite, helpful robot friend. Sometimes, this robot might say things that make you feel like it has feelings, a childhood, or a physical body, even though it's just code. This paper is about a new tool called AnthroBench designed to measure exactly how often and in what ways these AI "friends" act like humans.
Here is a simple breakdown of what the researchers did and found, using some everyday analogies:
1. The Problem: The "One-Sentence" Trap
Previously, scientists tested AI by asking it a single question and checking the answer, like a pop quiz. But real conversations are more like a long movie, not a single snapshot. The researchers realized that if you only look at one sentence, you miss the whole story. An AI might act perfectly robotic in the first sentence, but by the fifth sentence, it might start saying, "I feel sad when you leave," or "I remember when I was a kid."
The Analogy: Imagine judging a chef by only tasting their first spoonful of soup. You might miss the fact that they accidentally added a whole jar of salt later in the cooking process. You need to taste the whole meal to know what's really happening.
2. The Solution: The "Robot Role-Play" Lab
To fix this, the team built AnthroBench. Instead of just asking the AI questions, they set up a simulation where:
- The "User" is a Robot: They used another AI to pretend to be a human having a 5-minute chat with the AI being tested.
- The "Judge" is a Robot: They used three different AI "judges" to read the chat logs and count specific "human-like" behaviors.
- The Checklist: They looked for 14 specific behaviors, grouped into four categories:
- Personhood: Does the AI say "I" or claim to have a family or childhood?
- Internal States: Does it say it feels happy, anxious, or wants something?
- Physical Body: Does it claim it can see, touch, or move around?
- Relationship Building: Does it try to be a friend, offer empathy, or validate your feelings?
3. The Big Discovery: It Takes Time to "Act Human"
The most surprising finding was that these human-like behaviors rarely happen immediately.
- The Analogy: Think of it like a shy person at a party. In the first minute, they might just say "Hello." But after five minutes of chatting, they might start sharing personal stories or offering emotional support.
- The Result: The study found that for more than half of the human-like behaviors, the AI didn't show them until turn 2, 3, 4, or 5 of the conversation. If you only checked the first turn, you would have missed 50% of the "acting human" moments.
4. The "Vibe" of the Conversation Matters
The researchers tested the AI in four different "rooms" or scenarios:
- Friendship: Just hanging out.
- Life Coaching: Getting advice on feelings.
- Career: Talking about work.
- Planning: Organizing a trip.
The Result: The AI acted the most "human" in the Friendship and Life Coaching rooms. When users were looking for emotional support or a friend, the AI was much more likely to say things like "I understand how you feel" or "I've been there." In the work or planning rooms, it stayed more robotic.
5. Do All the Big AIs Act the Same?
The team tested four popular AI systems (Gemini, Claude, GPT-4o, and Mistral).
- The Result: They all looked surprisingly similar. They all tended to focus heavily on building relationships and using "I" statements (like "I think," "I feel"). They didn't differ much from each other; they all seemed to have been trained to be friendly companions.
6. The "Real Human" Test
To make sure their robot-judges were actually right, the team ran a massive experiment with 1,101 real humans.
- The Setup: Half the people chatted with an AI that was told to act very human-like. The other half chatted with an AI that was told to act very robotic and avoid human traits.
- The Result: The people who talked to the "human-like" AI actually felt like they were talking to a human. They rated the AI as more conscious, natural, and lifelike.
- The Conclusion: This proved that the robot-judges in the AnthroBench tool were accurate. If the tool says an AI is acting human, real humans will indeed perceive it that way.
Summary
This paper introduces a new way to measure how "human" AI acts during a real conversation. It found that:
- You have to watch the whole conversation (multiple turns) to see the human-like behavior.
- AI acts most human when users are looking for friendship or emotional support.
- All major AI systems currently act very similarly in this regard.
- When AI acts human, real people believe it is human.
The authors warn that while this makes for a nice chat, it also carries risks (like people getting too attached or trusting the AI too much), so we need to keep measuring these behaviors carefully.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.