Can LLMs Truly Embody Human Personality? Analyzing AI and Human Behavior Alignment in Dispute Resolution
This paper introduces an evaluation framework and dataset to investigate whether Large Language Models (LLMs) can accurately replicate human personality-driven behaviors in dispute resolution, ultimately finding that current LLMs diverge significantly from human patterns and may not be reliable proxies for human social behavior.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: Can AI Truly "Feel" the Vibe?
Imagine you are playing a video game where the characters are supposed to act like real people. You tell the game, "Make this character grumpy and stubborn," and you expect them to argue with you, refuse to compromise, and get defensive. You’d assume that because you gave them a "personality," they would behave like a real-life grumpy person.
This research paper asks a deep, skeptical question: When we tell an AI to act like a certain personality, is it actually "embodying" that personality, or is it just wearing a cheap mask that slips the moment things get heated?
To find out, the researchers put humans and AI into a "digital pressure cooker"—a simulated heated argument over a disputed online purchase—to see if their personalities actually changed how they fought and settled their differences.
The Experiment: The Digital Pressure Cooker
The researchers set up two different arenas:
- The Human Arena (KODIS): Real people having real, messy, emotional arguments about a disputed Kobe Bryant jersey.
- The AI Arena (L2L): Different AI models (like GPT-4, Claude, and Gemini) being "prompted" to act like people with specific Big Five personality traits (like being highly agreeable, neurotic, or extroverted).
They didn't just look at what was said; they looked at the strategy of the fight. They used a framework called IRP to categorize moves:
- The Peacekeeper (Cooperative): "I understand your frustration; let's find a middle ground."
- The Rule-Follower (Neutral): "Here are the facts of the transaction."
- The Aggressor (Competitive): "You're a liar! I'm not giving you a dime!"
The Findings: The Mask vs. The Soul
The results were a reality check for AI enthusiasts. The researchers found that while AI can mimic the words of a personality, it fails to capture the rhythm and soul of human behavior.
1. The "Transactional Robot" Problem (Strategy)
Think of a human argument like a jazz improvisation. Humans are unpredictable; they might start with facts, get emotional, try to compromise, and then suddenly get defensive. They adapt to the "vibe" of the person they are talking to.
AI, however, acts more like a pre-programmed vending machine. Even when told to be "agreeable," the AI tends to follow a very rigid, predictable pattern. It jumps straight to making deals or concessions without the natural "warm-up" or "grounding" phase that humans use. It’s "transactional" rather than "relational."
2. The "Emotional Mismatch" (Outcomes)
In the human world, personality is a massive driver of how a fight ends. For example, humans who are highly "neurotic" (emotionally sensitive) tend to react very differently to offers than those who aren't.
In the AI world, the "personality" didn't always translate to the outcome. While some AI models showed personality-driven results, they were often wrongly aligned. For instance, an AI might be told to be "agreeable," but it didn't actually show the same strategic patterns that a real-life agreeable human would use to reach a deal.
3. The "Rigid Actor" (Temporal Dynamics)
If you watch a movie, an actor's performance changes as the tension rises. Humans do this too—we change our tactics as an argument evolves.
The AI, however, is like an actor who reads the same line with the same tone from start to finish, regardless of whether the argument is just beginning or about to explode. They lack "temporal flexibility"—the ability to change their strategy as the "temperature" of the room rises.
Why Does This Matter? (The "So What?")
Imagine if we used these AI models to:
- Coach people on how to resolve conflicts.
- Act as mediators in legal disputes.
- Simulate social scenarios for training doctors or police officers.
If the AI is just wearing a "personality mask" and doesn't actually understand the deep, messy, psychological mechanics of how personality drives behavior, the training will be wrong. It would be like learning to drive by practicing in a simulator that doesn't account for gravity or friction.
The Bottom Line
The paper concludes that AI is a great mimic, but a poor actor. It can say the words, but it doesn't yet "live" the personality. Before we let AI handle high-stakes human social interactions, we need to move beyond giving it "adjectives" and start teaching it the deep psychological "why" behind human behavior.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.