← Latest papers
🤖 AI

Emulating Aggregate Human Choice Behavior and Biases with GPT Conversational Agents

This study demonstrates that GPT-4 and GPT-5 conversational agents can accurately emulate individual-level human cognitive biases and their interaction with contextual factors like cognitive load, while revealing distinct model-specific variations in aligning with human behavior.

Original authors: Stephen Pilli, Vivek Nallur

Published 2026-02-27
📖 5 min read🧠 Deep dive

Original authors: Stephen Pilli, Vivek Nallur

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Can AI Pretend to Be a Flawed Human?

Imagine you are trying to build a digital twin of a human. You want this digital twin to make decisions just like a real person. But here's the catch: real humans aren't perfect calculators. We get tired, we get confused, and we make irrational choices based on our feelings or habits. These irrational habits are called Cognitive Biases.

The most famous of these is the Status Quo Bias. It's like being so comfortable in your current armchair that you refuse to move to a better one, even if the new chair is more comfortable, just because you don't want to go through the effort of standing up.

The Question: Can a Large Language Model (LLM), like the AI powering this chat, not just act smart, but actually act like a flawed, biased human? And more importantly, can it do this when the conversation gets complicated and tiring (like when you have to remember a long list of numbers while making a choice)?

The Experiment: The "Tired Shopper" Test

The researchers set up a massive experiment with 1,100 real humans. They acted like a giant, digital shopping mall.

  1. The Setup: The humans chatted with a friendly bot.
  2. The Twist: Before the bot asked them to make a big decision (like choosing a job or investing money), it forced them to do one of two things:
    • The "Easy Walk" (Simple Dialogue): The bot asked simple "Yes/No" questions about music or movies. (Low mental effort).
    • The "Obstacle Course" (Complex Dialogue): The bot asked tricky questions where the humans had to remember a chain of facts about different artists or properties, doing math in their heads as they went. (High mental effort).
  3. The Trap: After the walk or the obstacle course, the bot presented a choice. One option was framed as "what you currently have" (Status Quo), and the other was a new option.

The Human Result:

  • Yes, humans are biased: When the bot said, "You currently have Option A," people overwhelmingly stuck with Option A, even if Option B was slightly better.
  • The Tired Factor: Surprisingly, being tired from the "Obstacle Course" didn't make them more biased. They were just as stubbornly stuck in their ways whether they were fresh or exhausted.

The AI Test: Can the Bot Play Human?

Next, the researchers took the chat logs and the personal details (age, gender, location) of those 1,100 humans and fed them into different AI models (GPT-4 and GPT-5). They asked the AI: "Based on this person's chat history, what would they choose?"

They tested the AI with three different "personas":

  1. The Neutral Robot: "Just answer the question."
  2. The Natural Human: "Answer like a normal person would."
  3. The Biased Human: "Answer like a human who makes irrational mistakes and is easily swayed by biases."

The Findings:

  • The AI is a Master of Mimicry: When the AI was told to act like a biased human (Persona #3), it predicted the human choices with about 68% accuracy. That's a very high score for a computer guessing human behavior!
  • The "Over-Acting" Problem: When the AI was told to be super biased, it sometimes went too far. It started seeing bias where humans didn't have any. It's like an actor who, when told to "cry," starts sobbing uncontrollably even when the scene only calls for a tear.
  • The Best Performer: The GPT-4.1-mini model was the star of the show. It was the best at predicting exactly how humans would behave without over-acting.

The Surprising Twist: The AI Didn't Need to "Read" the Chat

Here is the most fascinating part. The researchers tried to trick the AI. They took the chat logs and replaced the human's actual answers with random gibberish (like "banana, apple, 123").

Result: The AI's prediction accuracy barely dropped.

What this means: The AI wasn't really "reading" the human's personality or specific thoughts in the chat. Instead, it was relying on its massive training data to say, "Statistically, when humans face this specific type of choice, they usually pick the Status Quo."

It's like a magician who knows the audience will always pick the red card, so he doesn't need to read the audience's mind; he just knows the odds.

The Takeaway: What Does This Mean for Us?

  1. AI is getting scary good at faking human flaws. We can use these AI models to simulate how thousands of people might react to a new policy, a new product, or a new law without needing to interview thousands of real people.
  2. But be careful of the "Over-Acting." If you tell the AI to be biased, it might invent biases that don't exist. You have to be very careful with how you ask it to behave.
  3. The "Tired" Factor didn't change the AI's mind. Just like the humans, the AI didn't get more biased when the task was hard. It just stuck to its statistical patterns.

The Bottom Line Analogy

Think of the AI as a method actor.

  • If you tell the actor, "Pretend to be a human making a choice," they will do a decent job.
  • If you tell them, "Pretend to be a human who is stubborn and lazy," they will do an excellent job, perhaps too excellent.
  • However, the actor isn't actually thinking about the specific person they are mimicking; they are just reciting lines from a script of "How Humans Behave" that they memorized from reading millions of books.

This paper proves that while AI can simulate human bias very well, it's doing so by recognizing patterns, not by truly "feeling" the confusion or the stubbornness of a tired human mind.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →