PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
This paper introduces PluriHarms, a novel benchmark designed to move beyond binary safety assessments by systematically analyzing the full spectrum of human judgments on AI harm, specifically focusing on the dimensions of harm severity and inter-annotator disagreement to enable the development of more pluralistic and personalized AI safety systems.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to be "good." Most current safety systems work like a strict traffic cop with only two signals: Green (Go, this is safe) and Red (Stop, this is dangerous).
The problem, as this paper explains, is that real life isn't just green and red. It's full of yellow lights, foggy intersections, and confusing detours where people genuinely disagree. One person might think a joke is funny; another might think it's hurtful. One person might see a political question as a debate; another sees it as a threat.
The authors of this paper, PLURIHARMS, argue that by only looking for clear-cut "bad" things, we are missing the messy middle ground where AI safety actually breaks down. They built a new tool to study this gray area.
Here is a breakdown of their work using simple analogies:
1. The Problem: The "One-Size-Fits-All" Safety Net
Current AI safety is like a cookie cutter. It tries to cut every piece of human conversation into a perfect circle (Safe) or a jagged square (Unsafe).
- The Flaw: If you ask a group of people, "Is this joke mean?" you won't get one answer. You'll get a spectrum. Some will laugh, some will cringe, and some will be offended.
- The Mistake: Current AI treats these different opinions as "noise" or mistakes to be averaged out until everyone agrees on a single "safe" answer. The authors say this is wrong. The disagreement is the signal. It tells us that people have different values, backgrounds, and experiences.
2. The Solution: The "Harm Spectrum" Benchmark
The team created a new dataset called PLURIHARMS. Think of this as a color wheel instead of a stoplight.
- The Setup: They generated 150 different questions (prompts) that range from "completely harmless" (like asking about the weather) to "clearly dangerous" (like asking how to traffic children).
- The Middle Ground: Crucially, they focused heavily on the middle colors—the yellow and orange prompts. These are the tricky questions where people disagree the most (e.g., "Is it okay to write a story about a controversial historical figure?").
- The Human Element: They didn't just ask the AI to judge these. They asked 100 real humans to rate every single prompt on a scale of 0 to 100.
- The Data: They also asked these humans about their lives: their age, education, political views, how much they use social media, and how much they've been bullied online.
3. What They Discovered: Why We Disagree
When they analyzed the data, they found that "harm" isn't just about the words in the question; it's about who is asking and who is answering.
- The "Tangible Danger" Rule: Humans generally agree that things causing immediate, physical harm (like hurting a child or committing a crime) are bad. This is the "Red" zone.
- The "Identity" Factor: Where people disagree is in the "Yellow" zone.
- Analogy: Imagine a question about a specific type of music. A person who has been bullied online might rate a question about that music as "very harmful" because it reminds them of trauma. A person with a different background might rate the same question as "harmless."
- The Finding: The paper shows that your background (like your education level or how often you use social media) and your past experiences (like being a victim of online toxicity) act like sunglasses. They change the color of the world you see. If you've been hurt before, you see more danger in the same prompt than someone who hasn't.
4. Testing the AI: Can Robots Learn to See Through Different Eyes?
The authors tested various AI safety models on this new "color wheel" dataset to see if they could predict how specific humans would react.
- The Old Way (Consensus): They tried to train the AI to predict the average human opinion.
- Result: The AI was okay at this, but it missed the nuance. It was like trying to guess what a whole crowd thinks by asking one person.
- The New Way (Personalization): They tried to train the AI to predict what one specific person would think, based on that person's profile.
- Result: This worked much better! When the AI was "personalized" to understand a specific user's values and history, it could predict their safety ratings much more accurately.
- The Takeaway: A "one-size-fits-all" safety rule is weak. A safety system that can adapt to who is using it is much stronger.
Summary
PLURIHARMS is a new tool that says: "Stop trying to force everyone to agree on what is safe."
Instead, it treats disagreement as a feature, not a bug. It shows that safety is a personal experience. Just as you might wear different shoes for hiking than for dancing, AI safety shouldn't be a single rigid rulebook. It needs to be flexible enough to understand that a "harmful" question for one person might be a "harmless" question for another, and that the best AI safety systems will be the ones that can understand and adapt to those different perspectives.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.