← Latest papers
💬 NLP

Can LLMs Simulate Human Behavioral Variability? A Case Study in the Phonemic Fluency Task

This study demonstrates that while certain large language models can approximate human averages in the phonemic fluency task, they consistently fail to replicate the scope and structure of human behavioral variability, highlighting significant limitations in using LLMs as substitutes for human participants in cognitive research.

Original authors: Mengyang Qiu, Zoe Brisebois, Siena Sun

Published 2026-02-27
📖 5 min read🧠 Deep dive

Original authors: Mengyang Qiu, Zoe Brisebois, Siena Sun

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Question: Can AI Pretend to Be a Human?

Imagine you are a director casting a play. You need actors to play 106 different characters. Instead of hiring 106 real people, you decide to use a super-smart AI robot. You tell the robot, "Pretend to be a 35-year-old college grad. Now, say as many words as you can that start with the letter 'F' in one minute."

The big question this study asks is: Can a bunch of different AI robots, acting as different people, actually sound like a crowd of real, messy, unpredictable humans? Or will they all sound like the same robot wearing different hats?

The Experiment: The "F" Word Game

The researchers used a classic brain game called Phonemic Fluency.

  • The Task: You have 60 seconds. Say every word you can think of that starts with the letter "F." No proper names (like "Fred"), no numbers, and no repeating the same word with a different ending (like "fish," "fishing," "fished").
  • The Humans: They tested 106 real people.
  • The Robots: They tested 34 different Large Language Models (LLMs)—the brains behind chatbots like Claude, GPT, and Gemini. They even tried "thinking" modes (where the AI pauses to reason) and different settings.

They gave the robots the humans' age, education level, and how many words the humans actually said, hoping this would help the robots mimic the humans' performance.

The Results: The "Uncanny Valley" of Behavior

Here is what they found, broken down simply:

1. The Robots are Good at Math, Bad at Chaos

The AI models were surprisingly good at following the rules. Most of them managed to say roughly the same number of words as the humans. If a human said 15 words, the AI said 15 words.

  • The Analogy: Imagine a choir where every singer hits the exact same note perfectly. It sounds impressive, but it's not a real human choir. Real humans are messy; some sing 10 words, some sing 25, and some get stuck on "F" and only say 5. The AI was too perfect.

2. The "Variety" Problem

This is the biggest failure.

  • Humans: When you look at all 106 humans combined, they produced 476 unique words. Many people said rare, weird, or personal words (like "furbelow" or "fipple").
  • Robots: Even the best AI only produced about 226 unique words. They all stuck to the most common, boring words like "fun," "fire," "fan," and "family."
  • The Analogy: Imagine asking 100 people to draw a picture of a "dog."
    • Humans: You get 100 different drawings: a Chihuahua, a Golden Retriever, a cartoon dog, a dog in a hat, a dog made of spaghetti.
    • AI: You get 100 drawings that all look like the exact same Golden Retriever from a textbook. They are technically correct, but they lack the spark of individuality.

3. The "Thinking" Trap

The researchers thought, "Maybe if we tell the AI to think harder before answering, it will be more creative."

  • The Result: It did the opposite. When the AI used its "thinking" mode, it became even less diverse. It became more rigid and stuck to the safest, most common answers.
  • The Analogy: It's like asking a friend to tell a joke. If they just blurt it out, they might tell something weird and funny. If you tell them, "Stop and think really hard about the perfect joke," they will likely tell the safest, most generic joke they know so they don't mess up.

4. The "Ensemble" Failure (The Group Hug)

The researchers tried a clever trick. They thought, "If one robot is boring, maybe a group of robots will be interesting." They mixed the outputs of 33 different models, hoping the differences between the models would create a "super-human" variety.

  • The Result: It didn't work. The group was still boring.
  • The Analogy: Imagine you have 33 different chefs. You think, "If I mix their recipes, I'll get a huge variety of dishes!" But when you check their pantries, you realize they all bought their ingredients from the exact same warehouse. They all have the same flour, the same sugar, and the same spices. So, no matter how many chefs you mix, you're just getting the same basic cake over and over again. The AI models are all trained on the same internet data, so they all "know" the same common words.

5. How They "Think" is Different

The researchers looked at how the words were connected in the human brain versus the AI brain.

  • Humans: Our brains are like a messy city with tight neighborhoods. We have strong local connections (e.g., "fish" leads to "boat" leads to "water"), but the whole city isn't perfectly connected. We take weird detours.
  • AI: The AI's brain is like a perfectly efficient subway map. Everything is connected efficiently, but there are no weird detours. It takes the most direct, logical path every time.
  • The Takeaway: Humans retrieve words based on personal memories and weird associations. AI retrieves words based on statistical probability (what word is most likely to come next).

The Bottom Line

Can AI replace humans in psychological studies?
No, not yet.

While AI is great at sounding fluent and following instructions, it is terrible at simulating human messiness. It cannot replicate the unique, weird, and diverse ways real people think and speak.

  • For the future: Researchers shouldn't just swap humans for AI. Instead, they should use AI as a "baseline" (a standard, boring control) to measure how unique and flexible real humans actually are.

In short: AI is a perfect mimic of the average, but it has no soul for the outliers. And in human behavior, the outliers are often the most interesting part.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →