Prompts to Proxies: Emulating Human Preferences via a Compact LLM Ensemble
The paper introduces *Prompts to Proxies* (), a cost-effective, two-stage framework that emulates human preferences in social science research by constructing a diverse pool of LLM agents and aggregating them via L1-regularized regression to match target population data without requiring fine-tuning or sensitive demographic information.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you want to know what a whole country thinks about a specific topic, like "Should we build more parks?" or "How much do you trust scientists?" Usually, you'd have to call thousands of real people, which is expensive, slow, and sometimes impossible if you need to test a new idea quickly.
This paper introduces a clever way to use Artificial Intelligence (AI) to stand in for those real people. They call their method Prompts to Proxies (P2P).
Here is the simple breakdown of how it works, using some everyday analogies:
The Problem: The "Average" AI is Boring
If you just ask a standard AI chatbot, "What do Americans think about parks?", it will give you a very safe, middle-of-the-road answer. It's like asking one person to speak for a whole crowd; they usually just give the "average" opinion. But real crowds are messy! Some people love parks, some hate them, some don't care, and some think they are too expensive. A single AI chatbot tends to smooth all that diversity out, making it a bad stand-in for a real survey.
The Solution: The "Tasting Panel" Approach
Instead of trying to make one AI perfect, the authors built a system that creates a team of diverse AI characters and then mixes their answers together.
Think of it like a wine tasting panel.
- The Goal: You want to recreate the exact flavor profile of a specific vintage of wine (the real human population's opinion).
- The Mistake: Trying to find one single grape that tastes exactly like the whole bottle. (This is what other methods try to do, and it usually fails).
- The P2P Method: You gather a huge variety of different grapes (different AI characters). Some are sweet, some are sour, some are fruity, some are earthy. Then, you mix them in specific amounts to recreate that perfect vintage flavor.
How P2P Works (The Two-Stage Process)
The paper describes this as a two-step process:
Stage 1: Building the "Diverse Team" (Active Endowment Generation)
First, the system needs to create a big pool of different AI characters.
- The Trick: Instead of just asking the AI to "be random," the system uses a smart strategy called entropy-based adaptive sampling.
- The Analogy: Imagine a chef trying to create a menu that covers every possible taste. If they keep making dishes that everyone already likes (like plain bread), they aren't learning anything new. So, the chef looks at the dishes people didn't agree on (low entropy) and specifically asks the kitchen to invent new variations to cover those gaps.
- The Result: They end up with 300 different "personas" (e.g., a retired teacher from the Midwest, a young tech worker in the city, a farmer who worries about inflation). These aren't just random; they are carefully chosen to cover the widest possible range of opinions.
Stage 2: The "Mixologist" (Regression-Based Aggregation)
Now that they have a huge team of diverse AI characters, they need to figure out how much of each character to use to match the real human data.
- The Analogy: Imagine you have a recipe for a real soup (the actual survey results from real humans). You have a pantry full of 300 different ingredients (your AI characters). You don't need all 300. You need to find the perfect small handful of ingredients that, when mixed together, taste exactly like the real soup.
- The Math: The system uses a mathematical tool (L1-regularized regression) to act as the mixologist. It looks at the real survey data and says, "Okay, to get this result, we need 10% of the 'retired teacher' voice, 5% of the 'tech worker' voice, and 0% of the 'farmer' voice."
- The Magic: It does this without needing to know sensitive private details about real people (like their exact income or race). It just looks at the answers people gave and figures out which mix of AI voices matches those answers best.
Why This is a Big Deal
- It's Cheap and Fast: You don't need to train a massive, expensive AI model from scratch. You just use existing AI models and pay a tiny fee (about $0.80 per survey) to run the simulation.
- It's Accurate: When they tested this against real surveys from the US (American Trends Panel) and around the world (World Values Survey), their "AI mix" was much closer to real human opinions than just asking a standard AI or using other common methods.
- It Works with Little Data: Even if they only had a tiny bit of real human data to learn from (less than 3% of what other methods use), their system still performed very well.
The Bottom Line
The paper argues that you don't need one "super-AI" to understand human society. Instead, you need a diverse cast of AI actors, and a smart way to direct them so their combined performance sounds exactly like the real audience. This allows researchers to simulate public opinion quickly, cheaply, and accurately, without needing to interview thousands of real people or access their private data.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.