← Latest papers
💬 NLP

Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation

This paper resolves conflicting findings on using large language models for opinion simulation by distinguishing between "emulation" (generating individual responses) and "estimation" (predicting distributions), demonstrating that base models excel at the former while post-trained models are superior at the latter.

Original authors: Seth Grief-Albert, Jessica Bo, Difan Jiao, Ashton Anderson

Published 2026-08-05
📖 4 min read☕ Coffee break read

Original authors: Seth Grief-Albert, Jessica Bo, Difan Jiao, Ashton Anderson

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to understand the mood of a massive crowd without asking everyone individually. You might ask a few people what they think, or you might ask an expert to guess the overall vibe. This is the kind of puzzle scientists face when they try to use Artificial Intelligence (AI) to simulate human opinions. In the world of computer science, specifically in a field called "Natural Language Processing," researchers are building giant AI models that have read almost everything written on the internet. These models are incredibly good at predicting what words come next, which makes them seem like they could act as perfect stand-ins for real people. But here is the tricky part: just because an AI can write a convincing story doesn't mean it can accurately predict how a whole group of people will vote, feel about the economy, or answer a survey. Scientists are currently arguing over whether these AI "people" are actually good at mimicking human diversity or if they are just pretending to be something they aren't.

This paper steps into that argument to solve a mystery: why do some studies say AI is great at simulating human opinions, while others say it fails miserably? The authors, researchers from Queen's University and the University of Toronto, suggest that the confusion comes from mixing up two very different jobs. They call the first job Emulation. This is like hiring an actor to play a specific character. The AI generates a single response as if it were a real person, and if you ask enough "actors," their answers naturally add up to show the shape of the crowd's opinion. The second job is Estimation. This is like asking a statistician to look at a crowd and simply say, "I think 60% of them will choose option A." The researchers tested this idea by pitting two types of AI against each other: "Base" models (the raw, unpolished versions) and "Post-trained" models (the versions that have been fine-tuned to be helpful assistants).

The results of their experiments, run on real survey data from the Pew Research Center, reveal a clear split in talents. When the task was Emulation—generating individual text responses to see what the crowd looks like when you add them all up—the Base models were the winners. They produced responses that looked much more like real human data and did a better job of keeping the differences between groups (like Democrats vs. Republicans or rich vs. poor) distinct. In fact, the "Post-trained" models often suffered from what the authors call "persona collapse," where they all started sounding the same, like a choir of clones, losing the messy diversity of real people.

However, the story flips when the task is Estimation. When the researchers simply asked the AI to "predict the percentage of people who will choose this answer," the Post-trained models shined. They were much better at giving the correct numbers directly. The raw Base models were okay at this, but the polished, helpful assistants were significantly more accurate at guessing the distribution without having to generate a single sentence of fake text.

The authors suggest that this happens because the "Post-training" process teaches the AI to be a helpful assistant. When you ask an assistant to guess a crowd's opinion, it uses its knowledge to give a smart, direct answer. But when you ask that same assistant to act like a specific person, it gets stuck in its "helpful assistant" role, which makes it sound too uniform and lose the unique quirks of different demographics. The raw Base models, on the other hand, haven't been forced into that single "assistant" box, so they can draw from a wider, more chaotic pool of human voices when they are asked to generate text.

So, the paper concludes that there is no single "best" AI for simulating humans. If you need to generate text that feels like a diverse group of people (like for a video game or a social science simulation), you should use the raw Base models. But if you just need a quick, accurate prediction of what a group will think, you should ask the polished Post-trained models to estimate it. The confusion in the field wasn't because one type of AI was broken; it was because researchers were using the wrong tool for the specific job they were trying to do.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →