← Latest papers
💻 computer science

Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction

This paper introduces CoCoEval, a framework for detecting inconsistent and uncollaborative behaviors in LLM-simulated conversations, and finds that current models struggle to authentically replicate the complexity of human social interaction due to significant discrepancies in behavior frequency, the unreliability of prompt engineering, and the tendency of fine-tuning to produce narrow behavioral patterns.

Original authors: Ryo Kamoi, Ameya Godbole, Longqi Yang, Rui Zhang, Mengting Wan, Pei Zhou

Published 2026-03-19
📖 5 min read🧠 Deep dive

Original authors: Ryo Kamoi, Ameya Godbole, Longqi Yang, Rui Zhang, Mengting Wan, Pei Zhou

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to build a virtual movie set where AI actors (Large Language Models, or LLMs) play out human conversations. You want these AI actors to feel real, so you ask them to simulate a business meeting, a family argument, or a debate.

The researchers in this paper asked a simple but profound question: "Do these AI actors actually act like real humans, or are they just being too polite and perfect?"

Here is the breakdown of their findings, explained with some everyday analogies.

The Core Problem: The "Too Perfect" AI

Real human conversations are messy. We interrupt each other, we misunderstand what someone said, we repeat ourselves, we get distracted, and sometimes we stubbornly disagree even when the other person is right. We call these "inconsistent and uncollaborative behaviors."

However, when the researchers asked AI models to simulate these conversations, the AI actors were too well-behaved.

  • The Analogy: Imagine a group of friends playing a board game. In a real game, someone might accidentally knock over a piece, forget the rules, or argue about a move. In the AI version, everyone follows the rules perfectly, never interrupts, and always agrees instantly. It's like watching a robot dance where everyone moves in perfect, synchronized unison. It looks smooth, but it's boring and fake.

The New Tool: "COCOEVAL" (The Behavior Detective)

To prove this, the team built a new tool called COCOEVAL. Think of this as a behavioral detective that watches the AI conversations and counts specific "messy" human moments.

They defined 10 types of "messy" behaviors to look for, such as:

  1. The "Wait, what?" moment: Misunderstanding what someone just said.
  2. The "Stop talking!" moment: Interrupting someone mid-sentence.
  3. The "Broken Record" moment: Repeating the same thing over and over.
  4. The "Ghosting" moment: Answering a question with a vague "maybe" or changing the subject entirely.

What They Found (The Three Big Surprises)

The researchers tested the AI using three different methods, and the results were surprising:

1. The "Default" Setting (Vanilla Prompting)

When they just asked the AI, "Simulate a meeting," the AI produced conversations that were scarily clean.

  • The Result: The AI almost never interrupted, misunderstood, or disagreed. It was like a meeting where everyone is a mind-reader who never makes a mistake.
  • The Takeaway: By default, AI tries to be helpful and polite, which makes it terrible at simulating the messy reality of human interaction.

2. The "Instruction" Setting (Prompt Engineering)

The researchers tried to fix this by giving the AI a specific instruction: "Hey, please act messy! Interrupt people, misunderstand them, and disagree!"

  • The Result: The AI went too far. Instead of acting like a real human, it started acting like a chaotic cartoon character. It interrupted constantly and disagreed with everything, even when it didn't make sense.
  • The Analogy: It's like telling a shy actor, "Be loud!" They don't just become confident; they start screaming and knocking over chairs. The AI couldn't find the "sweet spot" of natural messiness; it just swung to the extreme.

3. The "Training" Setting (Supervised Fine-Tuning)

They tried training the AI on thousands of real human transcripts so it could "learn" how to be messy.

  • The Result: The AI learned to be messy, but only in one specific way. It became obsessed with repetition. It would say "Yes," then "Right," then "Exactly," over and over again, like a broken record.
  • The Takeaway: The AI learned to mimic the surface of human error (repeating words) but missed the depth (actual misunderstandings or genuine disagreements). It was like a student who memorized the word "oops" but didn't understand what a mistake actually feels like.

The "Turn" Factor

They also noticed that how the AI generates the conversation matters.

  • If the AI writes the whole conversation in one go (like writing a whole chapter at once), it stays very clean and logical.
  • If the AI writes just one sentence at a time (like a real person thinking before speaking), it gets messier. But even then, it struggles to get the right kind of messiness.

Why Does This Matter?

The authors warn us that if we use these AI simulations to study human behavior (like testing how people react to a new law or a social media feature), we might get the wrong answers.

  • The Metaphor: Imagine you are a city planner trying to design a traffic system. You build a simulation with robots that never run red lights, never get distracted, and always drive perfectly. You might think, "Great, traffic flows perfectly!" But when you put real humans in cars, the system crashes because real people get distracted, make mistakes, and cut each other off.

The Bottom Line

Simulating human conversation is much harder than it looks. It's not just about being smart or logical; it's about being flawed.

Currently, our AI actors are like perfectly rehearsed theater students who can't improvise a real, messy argument. They are either too polite, too chaotic, or just repetitive. Until we can teach them to be genuinely, naturally messy, we have to be very careful about trusting them to represent how real humans interact.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →