← Latest papers
🤖 AI

Using Zero-Shot LLM-Generated Survey Data for Geographically Explicit Population Synthesis

This paper evaluates the feasibility of using zero-shot LLM-generated health survey data as inputs for geographically explicit population synthesis, finding that while models like GPT-4.1 and Gemini-2.5-Pro can capture major state-level contrasts and produce reasonable tract-level patterns, their variable-dependent performance and the amplification of errors by iterative proportional fitting indicate they are currently suitable only as supplementary inputs rather than replacements for real survey data.

Original authors: Taylor Anderson, Sara Von Hoene, Orhan Yagizer Cinar, Emma Von Hoene, Amira Roess, Andrew Crooks, Hamdi Kavak

Published 2026-05-28
📖 4 min read☕ Coffee break read

Original authors: Taylor Anderson, Sara Von Hoene, Orhan Yagizer Cinar, Emma Von Hoene, Amira Roess, Andrew Crooks, Hamdi Kavak

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a city planner trying to build a perfect, miniature model of a real city to test how a new traffic light or a flu vaccine would work. To do this, you need a "synthetic population"—a digital crowd of millions of fake people who look, act, and live exactly like the real people in that city.

Usually, to build this crowd, you need a massive, real survey of actual people. But what if that survey doesn't exist, or is too expensive to get?

This paper asks a big question: Can we use a super-smart AI (like GPT-4 or Gemini) to just "dream up" these fake people from scratch, without showing it any real examples first? This is called "zero-shot" generation.

Here is how the researchers tested this, explained with simple analogies:

The Experiment: Two States, Two AIs

The researchers picked two very different U.S. states: Colorado (generally wealthier, more educated, and healthier) and Mississippi (with different demographic and health profiles).

They asked two AI models (GPT-4.1 and Gemini-2.5-Pro) to write survey answers for thousands of people in these states. Crucially, they didn't give the AI any real data to copy. They just said, "Imagine you are an adult in Colorado/Mississippi and answer these health questions."

The "Recipe" vs. The "Dish"

The researchers then took these AI-generated answers and fed them into a standard computer program called IPF (Iterative Proportional Fitting).

  • The Analogy: Think of the AI-generated survey data as a rough recipe written by a chef who has never tasted the real food. The IPF process is the cooking. The cooking process tries to adjust the ingredients so the final dish matches the known size of the city (the census data).
  • The Goal: They wanted to see if the "AI recipe" was good enough to cook a "digital city" that looked like the real one.

What They Found

1. The AI got the "Big Picture" right, but missed the details.
The AI models were surprisingly good at capturing the general vibe of the two states. They knew Colorado was generally healthier and wealthier than Mississippi.

  • The Metaphor: If you asked the AI to describe the weather in two different cities, it would correctly say, "City A is sunny and warm, City B is rainy and cool."
  • The Problem: When it came to specific details, the AI stumbled. It was great at guessing age or gender (like getting the color of the sky right), but terrible at guessing specific things like health insurance status or whether someone smokes (like getting the exact temperature wrong).

2. The "Cooking" Process (IPF) was a mixed bag.
Once the AI's rough recipe was fed into the cooking program, the results were unpredictable:

  • Sometimes it fixed the AI: For some variables, the cooking process smoothed out the AI's mistakes, making the final digital population look more like the real one.
  • Sometimes it made it worse: For other variables, the cooking process amplified the AI's errors. If the AI guessed wrong about how many people had insurance, the final model might be even more wrong about insurance than the original AI guess.

3. The "Map" Test.
The researchers checked if the fake people were living in the right neighborhoods (census tracts).

  • General Health: The AI did a decent job here. The digital map of "who is healthy" looked somewhat like the real map, especially in Mississippi.
  • Health Insurance: The AI failed miserably here. It consistently guessed that way too many people had insurance and too few were uninsured. This error stayed in the final map, making the digital city look much more insured than reality.

The Verdict

The paper concludes that using AI to generate survey data is promising but not ready for prime time.

  • The Good News: It's better than nothing. If you have no real data, an AI can give you a rough draft that captures the general differences between places.
  • The Bad News: It cannot replace real surveys yet. The AI makes specific mistakes that can mess up important policy decisions. You can't just swap a real survey for an AI one and expect the results to be perfect.

In short: The AI is like a talented artist who can paint a beautiful landscape from memory, but if you need to count the exact number of bricks in a wall to build a bridge, you still need to go measure the real wall yourself. The AI's painting is a great starting point, but you can't build the bridge on it alone.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →