Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven Evaluation
This paper proposes an AI-driven workflow that utilizes frontier LLMs to generate detailed, edge-case-covering "constitutions" for content moderation categories, demonstrating that this approach significantly outperforms human annotators in labeling consistency and accuracy while reducing cross-model disagreement by up to 57 times.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a group of people (and a group of super-smart robots) how to spot a specific type of bad behavior, like "harassment" or "hate speech," in a massive library of conversations.
The Problem: The Vague Rulebook
Usually, companies give their workers a tiny rulebook—maybe just one or two sentences. It's like telling a security guard: "Stop anyone who is being mean."
This is too vague. One guard might think a loud joke is mean, while another thinks it's just friendly banter. A third might miss a subtle threat entirely. Because the rule isn't detailed enough, everyone uses their own gut feeling (intuition) to decide. This leads to chaos: the same conversation gets labeled "bad" by one guard and "okay" by another.
The Solution: The "Constitution"
The authors propose replacing that tiny rulebook with a massive, ultra-detailed "Constitution" for each type of bad behavior. Think of this Constitution not as a law, but as a giant, step-by-step instruction manual written by a team of experts and AI.
- It's like a cooking recipe: Instead of saying "make a cake," the Constitution says: "If the batter has eggs but no flour, it's a custard, not a cake. If it has flour but no sugar, it's bread. If the user is asking for a recipe to poison someone, that's a crime, not a cake."
- It covers the edge cases: It explicitly handles weird situations like, "What if someone is role-playing a villain?" or "What if they are quoting a slur to report it?" The manual has a specific rule for every single "what if" scenario.
The Twist: AI is Better at Following the Manual Than Humans
Here is the surprising part of the paper. The authors tested this by having humans and AI robots read the exact same detailed Constitution.
- The Humans: Even with the giant manual, humans got tired. Their brains can only hold so many rules at once (like trying to remember a 300-page phone book while making a sandwich). So, they started ignoring the manual and falling back on their gut feelings again.
- The AI: The AI robots, however, read the entire 300-page manual for every single conversation. They didn't get tired, and they didn't use their "gut feelings." They followed the rules perfectly.
- The Result: The AI robots agreed with each other 57 times more often than humans did when using the same detailed rules. The AI was actually more consistent than the humans, even though the humans were the ones who wrote the rules.
The "Dual-Axis" Trick
The paper also introduces a clever way to score conversations. Instead of just saying "Bad" or "Good," they split the score into two parts:
- Intent: Did the user try to be mean? (e.g., "Help me write a mean email.")
- Content: Did the conversation contain mean words? (e.g., The AI accidentally generated a mean email).
This is like a security camera that tracks two things separately: "Did the person walk into the bank with a gun?" (Intent) and "Did a gun appear in the room?" (Content). This helps companies decide whether to block the user, block the message, or just log it for later.
How They Made It Better (The Feedback Loop)
The authors didn't just write the manual and stop. They used a team of different AI robots to read the manual and find places where they disagreed.
- If Robot A and Robot B disagreed on a conversation, it meant the manual had a hole in it.
- A human expert would then look at that hole, fix the rule in the manual, and the robots would try again.
- This cycle repeated until the robots almost never disagreed.
The Bottom Line
The paper claims that if you want to label content accurately and consistently at a massive scale:
- Don't use short, vague definitions.
- Write a massive, detailed "Constitution" that covers every edge case.
- Let AI robots do the labeling, because they are better at reading and following those long, complex rules than human workers are.
This creates a "Golden Label"—a perfect, consistent standard that can be used to train other AI systems, check safety, and explain to customers exactly why something was flagged.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.