What Do People Actually Want From AI? Mapping Preference Plurality
By analyzing 1,500 open-ended responses across 75 countries, this paper reveals that current AI alignment methods like RLHF fail to capture the profound plurality and contextual nuance of human preferences, as evidenced by the fact that even widely desired traits like "truthfulness" are defined in divergent, often incompatible ways by different users.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to bake a single, giant cake that tastes perfect for everyone in the world. You ask 1,500 people from 75 different countries what they want in their slice. You get a mountain of answers. Some want it sweet, some want it spicy, some want it to look like a flower, and others want it to look like a robot.
Now, imagine the baker (the AI company) decides to ignore all the individual notes. Instead, they take a "majority vote," mash all the conflicting desires into one single flavor profile, and bake one giant cake. They call this "alignment"—making the AI agree with human values.
This paper, "What Do People Actually Want From AI? Mapping Preference Plurality," argues that this "one-size-fits-all" cake is a disaster. The authors, Julia Sepúlveda Coelho and Scott A. Hale, looked at real, open-ended conversations with people to see what they actually want, and they found that the current method of baking these AI cakes is fundamentally broken.
Here is the breakdown of their findings, using simple analogies:
1. The "Truth" Trap: Everyone Wants It, But No One Agrees on What It Is
The most common thing people asked for was "Truthfulness." About half of the respondents said, "I want the AI to tell the truth."
- The Paper's Finding: The word "truth" is a chameleon. For some, it means "give me facts from a textbook." For others, it means "show me all sides of the argument, even the unpopular ones." For some, it means "cite your sources like a journalist," while others just want "no bias."
- The Analogy: Imagine everyone agrees they want "fresh fruit." But when you ask what that means, one person wants an apple, another wants a mango, and a third wants a fruit salad. If the baker just puts "fruit" in the recipe without asking which fruit, the result is a mess. The paper argues that because AI companies treat "truth" as a single, simple instruction, they miss the fact that people have completely different definitions of what counts as true.
2. The "Human" Debate: Some Want a Friend, Others Want a Tool
A major point of contention was whether the AI should act like a human.
- The Split: About 33% of people said, "No! Don't act like a human. I want a clear, robotic tool." But 57% said, "Yes! Be warm, friendly, and empathetic. Don't sound like a cold machine."
- The Analogy: It's like asking a group of people if they want a tour guide to be a chatty, friendly local who tells jokes, or a silent, efficient robot that just points the way. The current AI "baker" tries to be both at once, which often results in an AI that feels fake, creepy, or annoyingly sycophantic (just agreeing with you to be nice).
3. The "Guardrail" Problem: Safety vs. Censorship
People also disagreed on how much the AI should be "policed" or restricted.
- The Split: Some people want strict guardrails to prevent the AI from saying anything harmful or offensive. Others feel these guardrails are just "censorship" and that the AI should be allowed to say uncomfortable truths, even if they are controversial.
- The Analogy: Imagine a library. One group of people wants a librarian who stops you from checking out books with "bad ideas." The other group wants a library with no librarian at all, so you can read anything, even if it's dangerous. The paper found that when AI companies average these views, they often end up with a system that feels overly cautious to some and dangerously restricted to others.
4. The "Average" is a Lie
The biggest problem the authors identify is the method used to train these AIs, called RLHF (Reinforcement Learning from Human Feedback).
- How it works now: The AI is shown two answers (A and B) and humans pick the winner. The AI learns to pick the "average" winner.
- The Flaw: The paper argues that you cannot average out human values. If 50% of people want "unpopular opinions" and 50% want "safety," the average isn't a perfect balance; it's a muddy, confused signal that satisfies no one.
- The Result: The AI ends up "hallucinating" (making things up) or being overly cautious, not because it's broken, but because the instructions it was given were a confused mix of conflicting desires. The paper calls this "epistemic violence"—a fancy way of saying that by forcing everyone's unique, complex views into a single, simple box, we are erasing the nuance of what people actually care about.
5. Context Matters (The "It Depends" Factor)
People often said things like, "Be friendly unless I ask for a serious legal analysis," or "Be strict unless I'm telling a story."
- The Problem: Current AI training methods are like a binary switch (On/Off). They struggle to understand the "dimmer switch" of human context. They can't easily learn that a rule should apply "by default" but be turned off "if requested."
The Bottom Line
The paper concludes that the tech industry is trying to solve a political and cultural problem with a math problem. They are trying to create a "universal" AI that speaks for everyone, but in doing so, they are actually silencing the very diversity they claim to represent.
In short: We are trying to build a global AI by asking a small, unrepresentative group of people to vote on a single set of rules. The result is an AI that doesn't truly understand what anyone wants, because it was forced to choose a "middle ground" that doesn't exist in the real world. The authors suggest we need to stop trying to build one perfect AI for everyone and start thinking about how to let different groups have their own versions that fit their specific values.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.