← Latest papers
💬 NLP

The Enforcement and Feasibility of Hate Speech Moderation on Twitter

A global audit of Twitter reveals that despite the technical and economic feasibility of significantly reducing hate speech through human-AI moderation pipelines, the platform's current enforcement is largely ineffective due to institutional resource allocation choices rather than technical limitations.

Original authors: Manuel Tonneau, Dylan Thurgood, Diyi Liu, Niyati Malhotra, Victor Orozco-Olvera, Ralph Schroeder, Scott A. Hale, Manoel Horta Ribeiro, Paul Röttger, Samuel P. Fraiberger

Published 2026-04-15
📖 5 min read🧠 Deep dive

Original authors: Manuel Tonneau, Dylan Thurgood, Diyi Liu, Niyati Malhotra, Victor Orozco-Olvera, Ralph Schroeder, Scott A. Hale, Manoel Horta Ribeiro, Paul Röttger, Samuel P. Fraiberger

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine Twitter (now X) as a massive, bustling global town square where billions of people shout, whisper, argue, and share news every second. In this town, there's a serious problem: some people are spreading hate speech—verbal attacks that can hurt mental health and even incite real-world violence.

The city council (the platform) has a rulebook saying, "No hate speech allowed." They have a team of security guards (moderators) and a fleet of robot sentries (AI) to watch the square and remove bad actors.

But here's the big question this paper asks: Are the guards actually doing their job, and is it even possible for them to catch everyone?

The researchers decided to take a giant snapshot of this town square for a full 24 hours, looking at 375 million tweets. They then hired a huge team of human experts to read through 540,000 of these tweets to see which ones were actually hate speech. Five months later, they went back to see what happened to those specific tweets.

Here is what they found, broken down into simple stories:

1. The "Ghost Town" of Hate Speech

The Finding: Five months after posting, 80% of the hateful tweets were still online. Even the really scary, violent ones (like calls to kill specific groups) were mostly still there.

The Analogy: Imagine a town square where someone paints a swastika on a wall. You call the city council, and they say, "We'll get that." Five months later, you come back, and the paint is still there. In fact, the wall is just as dirty as it was day one. The study found that hateful tweets were no more likely to be removed than normal, boring tweets. It didn't matter if the tweet was violent or just mean; the platform treated them almost the same.

2. The Robot's Dilemma: Too Many False Alarms

The Finding: The researchers tested the "robot sentries" (AI). They found that if you tell the robot to catch every hate speech, it gets so confused that it starts arresting innocent people for just saying "stupid" or "jerk." If you tell the robot to be super careful so it doesn't arrest innocent people, it lets almost all the real hate speech walk right past.

The Analogy: Think of the AI like a metal detector at an airport.

  • Scenario A: You set the metal detector to be super sensitive. It beeps for everything—keys, belt buckles, even a soda can. You have to stop and search 10,000 innocent people to find 1 terrorist. This is too slow and annoying.
  • Scenario B: You set it to only beep for big guns. It misses the small knives and bombs hidden in shoes.
    The study found that the current robots are stuck in this middle ground. They can't perfectly tell the difference between "hate speech" and just "rude speech" without making a huge mess.

3. The Human Guard Shift: Not Enough Hands

The Finding: The researchers simulated a system where the robots act as a "triage nurse." The robot scans the tweets and says, "Hey, this one looks dangerous, let a human look at it." They found that even with the robots helping, the current number of human guards is way too low to catch the bad stuff.

The Analogy: Imagine the robot is a bouncer at a club who hands a "VIP list" to the human security team. The robot says, "These 100 people look sketchy." But the human team only has enough time to check 10 people before the night is over. The other 90 sketchy people slip right past the door. The study showed that the platform simply doesn't have enough human eyes to review the list the robots generate.

4. The "Money vs. Safety" Equation

The Finding: This is the most surprising part. The researchers calculated how much it would cost to hire enough humans to actually catch 80% of the hate speech that people see.

  • The Cost: It would cost about 2.9% of Twitter's total yearly revenue.
  • The Fine: If the government (like the EU or UK) catches them failing to stop this, they could fine the company 6% to 10% of their revenue.

The Analogy: Imagine a shop owner who knows their store is full of thieves.

  • Hiring enough security guards to stop the thieves would cost the owner $3.
  • But if the police catch them, the fine is $10.
  • Logic says: "I should just pay the $3 for guards."
  • Reality: The owner is still choosing not to hire the guards.

The study concludes that the reason hate speech isn't being stopped isn't because it's technically impossible or too expensive. It's a choice. The platform is choosing to spend less on safety than the cost of the fines they might get, or perhaps they are prioritizing other things (like keeping users engaged, even if that engagement comes from angry arguments).

The Big Takeaway

The persistence of online hate isn't a mystery of technology. It's not that the robots are too dumb or the humans are too slow. It's that the platform is making a business decision to allocate fewer resources to stopping hate speech than they would need to actually solve the problem.

They could fix it if they wanted to, and the cost is less than the fines they risk paying if they don't. But right now, they are letting the paint stay on the wall.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →