Synthetic Personalities: How Well Can LLMs Mimic Individual Respondents Using Socio-Economic Microdata?
This study demonstrates that large language models can effectively construct detailed individual-level "digital twins" from existing socio-economic panel data, achieving up to 78.8% accuracy and strong correlation with real respondents by optimizing information depth, embedding methods, and reasoning modes, thereby showing that twin-based market research is now limited by construction decisions rather than data availability.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to predict how a specific person would answer a survey question. Usually, you have to ask them directly, which is slow and expensive. This paper asks a bold question: Can we build a "digital clone" of a real person using only the data companies already have sitting in their databases?
Think of a company's database like a giant, messy attic filled with years of notes, receipts, and old surveys about their customers. The researchers wanted to see if they could take a random person's scattered notes from this attic and use an AI to create a "digital twin" that answers new questions just like the real person would.
Here is how they did it and what they found, explained simply:
1. The Ingredients: Building the Clone
To build these digital twins, the researchers used data from the German Socio-Economic Panel (SOEP). Imagine this as a massive, 40-year-long diary kept by over 28,000 real people. It covers everything from their jobs and income to their family life and opinions.
They picked 500 people from this diary and tried to build a digital version of each one. To do this, they tested a "recipe" with four main ingredients:
- The Brain (The AI): They used three different open-source AI models (think of them as three different types of super-smart calculators).
- The Memory (How much info to feed it): They tested feeding the AI just basic facts (like age and gender) versus feeding it the entire diary. They measured this "information depth" like filling a bucket with water.
- The Format (How to tell the story): They tried two ways to present the data:
- The Summary: A short, neat paragraph summarizing the person's life.
- The Raw Chat: A long, messy transcript of every question the person was ever asked and how they answered.
- The Thinking Style: They tested if the AI should just spit out an answer immediately, or if it should "think out loud" (write down its reasoning) before answering.
2. The Experiment: The "Taste Test"
The researchers created 2.1 million digital responses. They then asked these digital clones 183 questions that the real people had already answered in the past. The goal was to see if the digital clone's answer matched the real person's answer.
They measured success in two ways:
- Accuracy: Did the clone give the exact same answer? (e.g., If the real person said "Yes," did the clone say "Yes"?)
- Ranking: Did the clone understand who is more likely to say what? (e.g., If Person A is usually more optimistic than Person B, does the clone reflect that difference?)
3. The Big Discoveries
The "Goldilocks" Zone of Information
You might think "more data is always better." The researchers found this is only partly true.
- The First Sip: Adding a little more info (beyond just basic demographics) helped a lot.
- The Sweet Spot: The biggest jump in accuracy happened when they added the top 75% of the most interesting questions.
- Diminishing Returns: Adding the final 25% of the data (the last bit of the bucket) didn't help much. It cost a lot of computer power but barely improved the results.
- Analogy: It's like seasoning a soup. Adding salt helps a lot at first. Adding a little more helps. But dumping in the whole shaker at the end just makes it salty without making it taste much better.
The Power of Raw History
Surprisingly, feeding the AI the raw chat history (the messy transcript of past questions) worked better than giving it a neat summary.
- Analogy: It's like trying to guess what your friend will order for dinner. A summary says, "My friend likes Italian." But a chat history says, "Last Tuesday they ordered pizza, but when they were tired they got salad, and they hate cilantro." The messy history gave the AI more clues to get it right.
Thinking Helps, But Not for Accuracy
When the AI was told to "think out loud" before answering, it got better at ranking people (understanding the differences between individuals), but it didn't necessarily get the exact answers right more often.
- Analogy: It's like a student who writes out their math steps. They might get the final number wrong, but they understand the logic better than a student who just guesses the answer.
4. The Bottom Line
The study proves that you don't need to run special, expensive interviews to build good digital twins. You can build them using the heterogeneous, messy data companies already have (like loyalty program history or old surveys).
- Best Result: The most accurate digital twin got the right answer 78.8% of the time.
- The Takeaway: The biggest bottleneck for using these digital twins isn't how you design the data collection; it's just having enough data volume and picking the right AI model.
In short, if you have a digital attic full of customer notes, you can now build a surprisingly accurate "digital ghost" of your customers to test ideas, without needing to interview them one by one.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.