← Latest papers
🤖 machine learning

Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive

Original authors: Tharindu Cyril Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri, Christopher M. Homan, Ashiqur R. KhudaBukhsh

Published 2026-02-05
📖 5 min read🧠 Deep dive

Original authors: Tharindu Cyril Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri, Christopher M. Homan, Ashiqur R. KhudaBukhsh

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are the referee at a massive, chaotic sports game where the crowd is shouting comments at the players. Your job is to decide which comments are "offensive" and need to be silenced.

This paper is like a report card for two types of referees: Robots (computer algorithms) and Humans (real people with different political views). The researchers wanted to see if these referees agree on what counts as an insult, especially when the game is about American politics.

Here is the breakdown of their findings, using simple analogies:

1. The Robot Referees Can't Agree on the Rules

The researchers tested nine different computer programs (the "Machine Moderators") on over 3 million comments from YouTube news channels.

  • The Analogy: Imagine nine different referees watching the same play. One blows the whistle for a foul, another says "no foul," and a third says, "Well, it's a foul, but only if you're wearing blue shoes."
  • The Finding: The robots disagreed wildly. Sometimes, a comment was flagged as offensive by only one robot and ignored by the other eight. In fact, nearly half the comments were ignored by all the robots, while a tiny fraction was flagged by all of them. This means there is no single "robot truth" about what is offensive.

2. The Human Referees Have Different Playbooks

The researchers then asked real humans (Democrats, Republicans, and Independents) to rate the same comments. They asked two questions:

  1. Personal Offense: "Does this comment hurt your feelings?"
  2. Vicarious Offense: "Do you think this comment would hurt the feelings of someone from the other political team?"
  • The Analogy: It's like asking a fan of Team A, "Does this insult hurt you?" and then asking, "Do you think a fan of Team B would be hurt by this?"
  • The Finding:
    • Humans agree more with each other than with robots: Democrats, Republicans, and Independents generally understood each other better than the robots did.
    • The "Empathy Gap": However, humans were terrible at guessing what the other side would find offensive.
    • The Republican Blind Spot: The study found that Republicans were the worst at guessing what Democrats or Independents would find offensive. Conversely, Democrats and Independents also struggled to understand what would actually offend Republicans. It's like two groups of people shouting at each other, each convinced the other is shouting at the wrong volume, when in reality, they are shouting at completely different frequencies.

3. The "Vicarious Offense" Experiment

The paper introduces a new concept called Vicarious Offense.

  • The Concept: Usually, we only care if we are offended. This study asked: "Can you put yourself in someone else's shoes and predict if they would be offended?"
  • The Result: The researchers found that people are surprisingly bad at this. Even when Democrats and Independents agreed on what they thought would offend a Republican, they were often wrong about what actually did offend the Republican. They were all looking at the same comment but seeing different things.

4. The "Sensitive Topics" Trap

The researchers looked at specific hot-button issues like abortion and gun control.

  • The Analogy: If the game is about general sports, the referees might agree on the rules. But if the game is about a religious or deeply personal topic, the rules seem to change depending on who is holding the whistle.
  • The Finding: Disagreement got even worse on these topics.
    • On guns, Republicans were less tolerant of offensive comments than Democrats.
    • On abortion, Democrats were less tolerant than Republicans.
    • Independents were generally the most tolerant group, willing to let more comments slide than the two extreme sides.

5. The "Big Tech" Dilemma

The paper highlights a confusing reality for social media companies:

  • The Robots: They are inconsistent. One day they might delete a comment; the next day, a different robot might let it stay.
  • The Humans: They are biased by their politics. A Democrat might delete a comment a Republican thinks is fine, and vice versa.
  • The Conclusion: There is no perfect "Gold Standard" for what is offensive. Depending on who you ask (a robot, a Democrat, or a Republican), the answer changes completely.

Summary

The paper concludes that policing online speech is incredibly difficult because offense is in the eye of the beholder.

  • Robots are inconsistent and often miss the nuance.
  • Humans are biased by their political teams and are bad at guessing what the "other team" finds hurtful.
  • The Result: We are stuck in a situation where we cannot easily agree on what is "offensive," making it very hard to build fair systems to moderate online political arguments.

The authors also tested a fancy new AI (ChatGPT) to see if it could predict what the other side would find offensive, but it struggled just like the humans did, proving that this is a very hard problem for both brains and computers to solve.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →