Three Models of RLHF Annotation: Extension, Evidence, and Authority
This paper proposes distinguishing between three conceptual models of human annotators in RLHF—extension, evidence, and authority—to guide the design of more effective annotation pipelines by tailoring data collection and aggregation strategies to the specific normative role required for each dimension of model alignment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are building a very smart robot that can write stories, answer questions, and chat with people. You've taught it how to speak by feeding it millions of books, but now the robot is a bit wild. It sometimes says rude things, lies, or gives bad advice. To fix this, you hire a team of human "editors" to tell the robot what is good and what is bad. This process is called RLHF (Reinforcement Learning with Human Feedback).
The paper by Steve Coyne argues that when we hire these human editors, we often get confused about why we are hiring them. Are they there to copy our ideas? To tell us facts we don't know? Or to make the final law for us?
Coyne suggests there are three distinct ways to think about these human editors, and mixing them up causes the robot to get confused and behave poorly.
Here are the three models, explained with simple analogies:
1. The Extension Model: "The Shadow Clone"
The Idea: In this model, the human editors are there to copy the robot's designers.
The Analogy: Imagine you are a chef who wants to teach a sous-chef to cook your signature dish. You don't want the sous-chef to add their own unique twist or ask what the customers like. You want them to taste the food and say, "Yes, this tastes exactly like how I would have made it." If the sous-chef says, "I think it needs more salt," but you think it's perfect, you ignore them.
How it works: The editors are just extensions of the design team's brain. They are there to fill in the gaps when the designers are too busy to check every single answer.
- Good for: Deciding the robot's tone, style, or company-specific rules.
- The Trap: If you treat them like this but then get mad when they disagree with the "community," you are confused.
2. The Evidence Model: "The Expert Witness"
The Idea: In this model, the editors are there to provide facts or data that the designers don't have.
The Analogy: Imagine you are a judge in a courtroom, but you don't know anything about a specific type of rare fish. You hire a marine biologist as an expert witness. You ask, "Is this fish poisonous?" The biologist says, "Yes, based on my knowledge, it is." You must listen to them, even if you personally think it looks safe. Their job is to give you evidence about the truth.
How it works: The editors are experts or representatives of the public who tell the designers, "Actually, most people find this offensive," or "This fact is wrong." The designers shouldn't overrule them just because they disagree; the editors are providing data.
- Good for: Checking if the robot is factually accurate or if it violates general social standards (like "is this offensive?").
- The Trap: If you treat them like experts but then fire them because they don't agree with your personal opinion, you lose the value of their expertise.
3. The Authority Model: "The Mini-Legislature"
The Idea: In this model, the editors have the power to make the rules, regardless of whether they are "right" or "wrong" in a factual sense.
The Analogy: Imagine a town meeting where a group of randomly selected citizens votes on a new law. They don't need to be experts on traffic; they just need to represent the people. If they vote to lower the speed limit, the city must obey, not because the citizens are "smarter" than the mayor, but because they have the authority to decide for the community.
How it works: The editors act as a "mini-legislature." Their job isn't to find the truth or copy the designer; it's to exercise the right to decide what the robot should do. Their power comes from being a fair, representative sample of the population.
- Good for: Deciding on controversial political issues or deep moral questions where there is no single "correct" answer, only what the community decides.
- The Trap: If you treat them like a legislature but then ignore their vote because it doesn't match your personal view, you are breaking the "social contract" of the system.
Why Does Mixing Them Up Break the Robot?
The paper argues that most current AI systems are a messy mix of all three, which leads to failure modes:
- Self-Defeat: Imagine you hire editors to act as a "Mini-Legislature" (Authority) to represent the public. But then, you fire any editor who disagrees with your personal opinion (Extension). You have destroyed the very thing you asked for: a representative voice. You can't have a democracy where the leader can fire voters who vote against them.
- Fragmentation: Imagine you tell the editors, "You are here to give us facts" (Evidence), but your instructions are vague, so they think you just want them to copy your style (Extension). The result is a pile of confusing data where some editors are trying to be experts and others are trying to be clones. The robot gets a mixed signal and learns nothing useful.
- Misattribution (The "Responsibility Laundering"): This is when a company says, "We let the public decide what the robot says!" (Authority), but in reality, they secretly fire anyone who disagrees with them (Extension). They are trying to shift the blame for bad robot behavior onto the "public" while actually keeping all the control.
The Author's Big Recommendation
Steve Coyne suggests we stop trying to build one single pipeline to do everything. Instead, we should build different pipelines for different jobs:
- For Style: Use the Extension model. Let the designers decide if the robot should sound excited or serious.
- For Facts: Use the Evidence model. Hire experts to tell the robot what is true.
- For Controversial Values: Use the Authority model. Let a representative group of people vote on what the robot should do in tricky moral situations.
The Bottom Line:
Building a helpful AI isn't just a technical problem; it's a philosophical one. Before you hire your human editors, you need to decide: Are they your clones, your experts, or your lawmakers? If you don't pick one clearly, your robot will end up confused, inconsistent, and potentially unfair.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.