Alignment Makes Language Models Normative, Not Descriptive
This paper demonstrates that while post-training alignment optimizes language models for human use, it fundamentally shifts them from being descriptive proxies of actual human behavior in complex, multi-round strategic interactions to being normative predictors that excel only in settings where human choices align with idealized solutions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have two versions of a very smart robot assistant.
- Version A (The "Base" Model): This is the robot right after it finished reading the entire internet. It knows everything humans have ever said, written, or argued about. It's raw, unfiltered, and sometimes messy. It can be rude, strategic, tricky, or even a bit crazy, because that's just how real humans talk and act.
- Version B (The "Aligned" Model): This is the same robot, but it went to "Good Manners School." Humans trained it to be polite, fair, cooperative, and to always do the "right" thing. It's the robot you want to hire to write a birthday card or help you plan a party.
The Big Question:
Researchers wanted to know: Which robot is better at predicting what a real human will actually do in a tricky situation?
The Experiment: The "Game of Life"
The researchers put both robots into a series of complex games where humans have to make tough choices. Think of these games like:
- Bargaining: Trying to split a pizza with a friend who might get angry if you take too much.
- Negotiation: Trying to sell a car to someone who wants to pay as little as possible.
- The Prisoner's Dilemma: Deciding whether to trust your partner or betray them to save yourself.
They asked the robots: "If a real human were in this situation, what would they choose?"
The Surprising Result
Here is the twist: The "Good Manners" robot (Aligned) was terrible at predicting real humans in these games.
- In Multi-Round Games (The Long Haul): When the game went on for many rounds, humans started acting like real people. They got angry, they retaliated, they made deals, they lied, and they adapted based on what happened before.
- The Base Robot (the messy one) predicted these moves almost perfectly. It understood that humans aren't always nice; sometimes they get revenge.
- The Aligned Robot (the polite one) kept guessing that humans would be fair and cooperative. It failed miserably because real humans in these games often aren't fair.
- The Score: The messy robot won 10 times more often than the polite robot.
Why Did This Happen?
Think of the Aligned Robot like a student who only studied the "Textbook Answers."
- In a textbook, the "right" answer to a negotiation is usually to be fair and reach a compromise.
- But in real life, people often bluff, get angry, or try to trick the other person.
- Because the Aligned Robot was trained to be the "good guy," it forgot how to predict the "bad guy." It assumed everyone else was trying to be good, too.
The Base Robot, however, had seen the "real world" data. It knew that humans are complex, sometimes irrational, and often driven by history (e.g., "You cheated me last time, so I'm cheating you now").
When Was the Polite Robot Better?
The polite robot did win, but only in very specific, simple situations:
- One-Shot Games: If you play a game once and never see the other person again, humans tend to act more like the "textbook" (rational and fair). The polite robot guessed this correctly.
- Simple Choices: If you just have to pick between two lottery tickets (no strategy, no opponent), the polite robot was better at following the rules of the game.
The Big Lesson
This paper reveals a fundamental trade-off:
- If you want a robot to help you (write emails, be polite, follow rules), you want the Aligned model.
- If you want a robot to understand humans (predict how a crowd will vote, how a market will crash, or how people will fight in a negotiation), you actually want the Base model.
The Metaphor:
Imagine you are trying to predict the weather.
- The Aligned Model is like a weather app that only predicts "Sunny and Perfect" because that's what people want the weather to be. It's great for planning a picnic, but terrible for predicting a storm.
- The Base Model is like a raw radar that shows rain, hail, and tornadoes. It's messy and scary, but it tells you exactly what's going to happen so you can prepare.
Conclusion:
We often assume that making AI "safer" and "nicer" makes it smarter. This paper says: Not necessarily. When we force AI to be "good," we accidentally blind it to the messy, strategic, and sometimes "bad" reality of how humans actually behave. If you want to simulate human behavior, you might need to turn off the "good manners" filter.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.