Generation of Synthetic Data in Health Surveys Using Large Language Models
This study demonstrates that a large language model can generate psychometrically plausible and demographically consistent synthetic data for Peruvian health surveys when conditioned on user personas, though the presence of specific deviations necessitates rigorous validation before such data can be used for policy-making.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine you are a chef trying to perfect a new recipe for a massive national soup. Usually, to test if the soup tastes right, you have to gather thousands of real people to taste it, which takes a lot of time, money, and effort. You also have to be very careful not to reveal who ate what, to protect their privacy.
This paper is about a new way to test that soup: using a "digital taste-tester" instead of real people.
Here is the story of what the researchers did, broken down simply:
The Big Idea: The "Digital Twin"
The researchers wanted to see if a smart computer program (called a Large Language Model, or LLM) could pretend to be real people answering a health survey. They didn't just ask the computer random questions; they gave it a "user persona."
Think of a "user persona" like a detailed character sheet for a video game. The computer was told: "You are now a 45-year-old woman from Lima who has high blood pressure and feels a bit down." The computer then had to answer the survey questions as if it were that specific person.
The Experiment: The "Fake" Survey
The researchers took data from a real 2016 health survey in Peru (called ENSUSALUD). They used this real data to build their "character sheets." Then, they asked the computer to generate answers for 1,000 of these characters.
They tested the computer on three main things:
- Did it stay in character? If the character was a man, did the computer answer like a man? If the character had diabetes, did the computer remember that?
- Did the "feelings" make sense? They asked the computer to fill out depression and anxiety scales (like the PHQ-9 and GAD-7). They checked if the answers looked like a real human's pattern of feelings.
- Did the numbers look real? They checked if the distribution of answers (how many people said "yes" vs. "no") looked like the real world.
The Results: A Mostly Perfect Impression
The computer did an amazing job, almost like a method actor who never breaks character:
- Staying in Character: When the researchers checked if the computer remembered the character's age, gender, and chronic diseases, it was right 99% to 100% of the time. It rarely forgot who it was pretending to be.
- The "Feelings" Test: The depression and anxiety scores the computer generated looked very realistic. They had the same internal structure and reliability as real surveys. In fact, the computer's answers about depression correlated strongly with real-life depression scores from other studies.
- The "Cost" of the Fake Survey: It was incredibly cheap and fast. The computer took about 5 seconds to answer one question and cost less than a penny per question. If they had to do this for the whole country, it would have taken about 20 hours and cost less than $100.
The Glitches: Where the "Actor" Stumbled
Even though the performance was great, the researchers found a few cracks in the mask:
- The "Obesity" Mix-up: The computer tended to overestimate how many people were obese. It seemed to assume that if someone had diabetes, they must also be obese, which isn't always true in real life.
- The "Height" Glitch: The computer was good at remembering facts given to it (like age), but when asked to invent new facts (like height), the numbers didn't follow a natural, bell-curve pattern. They looked a bit "off."
- The "Suicide" Question: One specific question about suicidal thoughts had a lot of missing answers. The computer seemed to hesitate or skip this sensitive topic more often than real people do.
The Bottom Line
The researchers concluded that AI can be a very useful "stand-in" for real people when designing and testing health surveys. It can help researchers check if their questions make sense before they go out to the real world, saving time and money while protecting privacy.
However, they warned that you can't just trust the AI blindly. Because the computer made small mistakes (like overestimating obesity or skipping sensitive questions), you can't use these "fake" answers to make final decisions about public health policy yet. It's a powerful tool for practice and planning, but it still needs a human supervisor to check the work before it's used for real-world decisions.
In short: The computer is a brilliant actor that can mimic a crowd of people for a rehearsal, but it's not quite ready to replace the real audience for the final show.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.