Who and What? Using Linguistic Features and Annotator Characteristics to Analyze Annotation Variation
This paper presents the first large-scale analysis of four harmful language detection datasets to demonstrate that interactions between annotator characteristics and linguistic properties are crucial for understanding annotation variation, while revealing that effect patterns vary significantly across datasets, cautioning against broad generalizations.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hosting a massive dinner party where 25,000 guests (the text items) are being judged by over 8,000 different food critics (the annotators). The goal is to decide which dishes are "toxic" or "offensive."
In the past, researchers tried to solve this by asking, "What is the average score?" and creating one single "Gold Standard" label for every dish. But this paper argues that approach misses the point. It's like asking, "Is this soup too salty?" and ignoring the fact that one critic is a salt-lover, another is on a low-sodium diet, and a third just had a bad day.
The authors of this paper ask a different question: "Who thinks what, and why?" They want to understand the complex dance between the person doing the judging and the thing being judged.
Here is the breakdown of their findings using simple analogies:
1. The Setup: A Cross-Check of Perspectives
The researchers looked at four huge datasets of online comments (like Reddit or Twitter) where people rated how offensive the text was.
- The "Who": They didn't just look at basic stats like age or gender. They looked at attitudes (e.g., "Do you think hate speech is a problem?") and moral beliefs.
- The "What": They analyzed the text itself, looking at everything from the number of swear words to the complexity of the sentences and the presence of specific topics (like politics or religion).
- The Twist: They didn't just look at these factors separately. They looked at how they interact. It's not just "Old people hate this" or "This text is bad." It's "Old people hate this specific type of text, but young people don't."
2. The Main Discovery: It's About the Mix, Not Just the Ingredients
The biggest takeaway is that you cannot predict how someone will react to a text just by knowing who they are, and you cannot predict how a text will be received just by reading it. The reaction happens in the collision between the two.
The "Age vs. Hate Words" Analogy:
In one dataset, the researchers found that for younger annotators, the presence of hateful words didn't change their rating much. But for older annotators, the presence of those same words made a huge difference. It's like a volume knob: for some people, the "offensiveness" of a text is turned up high only when they are older; for others, the volume stays the same regardless of age.The "Moral Compass" Analogy:
In another dataset, they found that people who care deeply about "protecting the vulnerable" (a moral value) rated long, complex texts differently than those who didn't. It wasn't just the length of the text; it was how that length interacted with their specific moral concern.
3. The "Hidden Traps" (Spurious Signals)
The study found some weird, accidental patterns that could trick a computer.
- The "Hindu Annotator" Glitch: In one dataset, they noticed that items containing strange slang or misspellings (non-standard words) were consistently rated as "not hateful" by annotators who identified as Hindu.
- The Lesson: This wasn't because Hindu people have a special immunity to hate speech. It was likely a coincidence in how the data was collected (maybe those specific weird texts happened to be shown to that group). If a computer learned this, it would think, "Oh, if it has a typo, it's safe!" which would be a dangerous mistake.
4. The "Batch" Problem
The researchers also simulated what happens if you swap out the group of judges.
- The Result: Even if you have the same text and the same guidelines, if you change the specific group of people doing the judging, the "winning" patterns change.
- The Metaphor: Imagine you ask Group A to judge a movie. They love the action scenes. You ask Group B (who look similar on paper) to judge the same movie. They might hate the action scenes and love the dialogue. You can't assume Group B will act exactly like Group A just because they are both "movie lovers."
5. What This Means for the Future (According to the Paper)
The authors don't claim this will immediately fix the internet or cure hate speech. Instead, they offer a few practical warnings for anyone building AI systems:
- Don't Flatten the Data: Don't just average out all the opinions into one number. That throws away the most interesting information. The disagreement is the signal.
- Be Careful with Generalizations: Just because a rule works for one group of people on one set of texts, it doesn't mean it works for everyone. A model trained on one group might fail miserably on another.
- Collect Better Data: If you want to understand toxic language, you need to know who is labeling it. You need to know their attitudes and backgrounds, not just their age.
- Watch for "Ghost" Patterns: Be careful that your AI isn't learning accidental patterns (like the "Hindu annotator" glitch) instead of real rules about language.
Summary
Think of this paper as a map that says: "The world of human judgment is messy, interactive, and highly specific." You can't just build a machine that says "This is bad." You have to build a machine that understands who is saying it, what they are looking at, and how those two things bump into each other. If you ignore the "Who" and the "What" interacting, your model will be blind to the most important parts of human communication.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.