← Latest papers
💻 computer science

Six Fallacies in Substituting Large Language Models for Human Participants

This paper argues that large language models should not replace human participants in behavioral and psychological research due to six critical interpretive fallacies, advocating instead for their responsible use as pragmatic simulation tools that complement, rather than substitute, human evidence.

Original authors: Zhicheng Lin

Published 2026-06-26
📖 6 min read🧠 Deep dive

Original authors: Zhicheng Lin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a psychologist trying to understand how people think, feel, and make decisions. Traditionally, you'd need to interview real humans, which is slow, expensive, and tiring. Recently, a new idea has popped up: "Why not just ask a super-smart AI (a Large Language Model, or LLM) instead? It talks like a human, so maybe it is a human participant."

The paper you shared, "Six Fallacies in Substituting Large Language Models for Human Participants," argues that this idea is a trap. The author, Zhicheng Lin, says that while AI is impressive, treating it as a replacement for real people is like trying to study the ocean by looking at a high-resolution painting of it. The painting looks real, but it has no water, no waves, and no depth.

Here are the six main mistakes (fallacies) the paper identifies, explained with simple analogies:

1. The "Magic Trick" Mistake (Token Prediction vs. Intelligence)

The Fallacy: Thinking that because an AI can predict the next word in a sentence perfectly, it actually understands what it's saying.
The Analogy: Imagine a parrot that has memorized every line from a Shakespeare play. If you ask it, "What happens next?" it can recite the next line perfectly. But the parrot doesn't know what "love," "betrayal," or "death" actually feel like. It has no body, no eyes, and no heart. It just knows that the word "love" usually follows the word "true."
The Paper's Point: AI is a statistical parrot. It predicts text based on patterns, not because it has a mind or feelings. It has never felt the sun on its skin or tasted an apple. Therefore, its "intelligence" is different from human intelligence.

2. The "Average Joe" Mistake (The Average Human Fallacy)

The Fallacy: Assuming the AI represents the "average" human because it was trained on human text.
The Analogy: Imagine you want to know what the average person thinks about food. You ask a robot that has read every cookbook and food blog ever written. The robot will give you a very polite, perfectly balanced, and "correct" answer. But real people are messy! Some people hate vegetables, some eat weird combinations, and some lie about what they eat. The robot is like a "perfectly curated" human who never makes a mistake or has a weird quirk.
The Paper's Point: AI is trained on specific types of text (mostly Western, educated, internet-heavy) and then tweaked to be "helpful" and "safe." This makes it too perfect and too neutral. It doesn't capture the messy, biased, and diverse reality of real human populations.

3. The "Same Result, Different Engine" Mistake (Alignment as Explanation)

The Fallacy: Thinking that because the AI gives the same answer as a human, it must be thinking the same way.
The Analogy: Imagine a digital clock and a sundial both show it is 12:00 PM. They are "aligned." But the digital clock uses batteries and circuits, while the sundial uses the sun and shadows. If you study the sundial to understand how a battery works, you will fail. Just because the AI and a human give the same answer doesn't mean they used the same mental process to get there.
The Paper's Point: AI might get the right answer for the wrong reasons (like guessing patterns). We cannot assume the AI is "thinking" like a human just because the output looks the same.

4. The "Talking to a Ghost" Mistake (Anthropomorphism)

The Fallacy: Attributing feelings, beliefs, or a "soul" to the AI because it sounds so human.
The Analogy: When you talk to a very realistic ventriloquist dummy, you might feel like it has a personality. But if you ask the dummy, "Do you believe in ghosts?" and it says "Yes," it's not because it has a belief. It's because the dummy is programmed to say "Yes" in that context. The AI is the dummy; it has no inner life, no fears, and no desires.
The Paper's Point: When researchers say the AI "believes" or "feels," they are making a mistake. The AI is just a tool generating text. It has no consciousness.

5. The "One-Size-Fits-All" Mistake (Identity Essentialization)

The Fallacy: Thinking you can just type "Act as a Black woman" or "Act as a Chinese man" and get a perfect, realistic representation of that group.
The Analogy: Imagine asking a chef to "cook a generic Italian meal." They might make a pizza. But "Italian" isn't just one thing; it's a million different regions, families, and personal tastes. If you ask the AI to be a "Black woman," it might just give you a stereotype (a caricature) rather than the complex, unique reality of a real person who happens to be a Black woman.
The Paper's Point: Real human identity is fluid and complex. Reducing it to a simple label in a prompt creates a cartoon version of a person, not a real one.

6. The "Time-Traveler" Mistake (The Substitution Fallacy)

The Fallacy: Thinking AI data can replace real human data to study current events.
The Analogy: Imagine trying to understand what people are wearing today by looking at a photo album from 2020. The album is static; it doesn't change. If a new fashion trend starts tomorrow, the photo album won't know about it. AI is like that photo album. It is trained on data that stops at a certain date. It cannot adapt to new social movements or changing opinions in real-time.
The Paper's Point: AI is stuck in the past. If you use it to study current human behavior, you are studying a snapshot of history, not the living, breathing present.

The Bottom Line: What Should We Do?

The paper doesn't say AI is useless. It says we need to change how we use it.

  • Don't use AI as a replacement for people. You can't swap out a real human participant for a robot and expect the same scientific truth.
  • Do use AI as a "Simulation Tool." Think of AI like a flight simulator. Pilots use simulators to practice, test new ideas, and see how a plane might react. But you wouldn't say the simulator is the plane, and you wouldn't fly a real passenger in a simulator.
  • The Best Approach: Use AI to generate ideas or test hypotheses quickly (the simulator), but then always check your findings with real human data (the real flight) to make sure you are right.

In short: AI is a powerful mirror that reflects human language, but it is not a window into the human mind. We must be careful not to confuse the reflection with the reality.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →