← Latest papers
💬 NLP

IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language

This paper demonstrates that automated content moderation tools fail to capture the heterogeneous and subjective attitudes of marginalized communities toward reclaimed slurs, leading to the misclassification of solidarity-driven language as hate speech due to a lack of contextual understanding and significant disagreement even among in-group members.

Original authors: Christina Chance, Rebecca Pattichis, Arjun Subramonian, James He, Shruti Narayanan, Saadia Gabriel, Kai-Wei Chang

Published 2026-04-21
📖 5 min read🧠 Deep dive

Original authors: Christina Chance, Rebecca Pattichis, Arjun Subramonian, James He, Shruti Narayanan, Saadia Gabriel, Kai-Wei Chang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: The "One-Size-Fits-All" Robot vs. The Complex Human Heart

Imagine you are trying to teach a robot how to understand human language. You give the robot a list of "bad words" (slurs) and tell it: "If you see this word, delete the post immediately because it's hate speech."

The robot does exactly what you ask. It deletes everything containing those words.

But here is the problem: The robot doesn't understand context, history, or friendship.

For many marginalized communities (like Black people, LGBTQIA+ people, and women), these same "bad words" have been taken back and used as a way to show love, solidarity, and pride among friends. It's like a group of friends using a nickname that would be an insult if a stranger used it, but a term of endearment if a best friend uses it.

This paper argues that current AI moderation tools are like that clumsy robot. They are too dumb to tell the difference between a friend joking and a stranger attacking. Because they can't tell the difference, they accidentally silence the very people they are supposed to protect.


The Core Experiment: Asking the Community to Judge

The researchers didn't just guess; they asked members of these communities to act as judges. They gathered about 12,000 tweets containing reclaimed slurs (the "n-word," the "f-word," and the "b-word") and asked community members to label them.

They asked questions like:

  • "Is this person using the word to be mean, or to show pride?"
  • "If the author is part of our community, should we report this?"
  • "If the author is not part of our community, should we report this?"

The Shocking Discovery: "We Don't All Agree"

The researchers expected that if they asked members of the same community, they would all agree. They thought, "If we ask Black people about the n-word, they will all say the same thing."

They were wrong.

The Analogy: Imagine a potluck dinner where everyone brings a dish. You expect everyone to agree on the recipe. Instead, you find that some people think the dish is a delicious family tradition, others think it's a funny inside joke, and some think it's actually gross.

The study found massive disagreement even among people from the same community.

  • One person might see a tweet and think, "That's a proud reclaiming of our identity!"
  • Another person from the same community might see the exact same tweet and think, "That's still hurtful and should be banned."

Why? Because everyone has a different life story.

  • Maybe one person grew up hearing the word used only as a weapon, so they can't stand to see it used anywhere.
  • Maybe another person grew up hearing it used as a joke among cousins, so they feel it's safe to use.

The paper calls this "Heterogeneous Attitudes." It means: There is no single "correct" way to feel about these words, even within the same group.

The Robot vs. The Humans

The researchers also tested a popular AI tool called Perspective API (which many social media sites use) against the human judges.

  • The AI's Logic: "I see the word. It is a slur. Therefore, it is hate speech. Delete it."
  • The Humans' Logic: "I see the word. Who said it? To whom? In what context? Is it a joke? Is it a quote? Is it a new slang meaning?"

The Result: The AI was often wrong.

  • It frequently flagged friendly, reclaimed usage as hate speech (False Positives).
  • It failed to understand that the same word might be acceptable if a Black person says it, but unacceptable if a white person says it.
  • The AI seemed to assume that everyone using the word was an outsider trying to be mean.

The "B-Word" vs. The "N-Word" Analogy

The paper found that different words are treated differently, even by the same people.

  • The "B-Word" (B*tch): This word has become so common in movies, music, and marketing that it's almost like a regular noun. The AI and humans agreed more on this one. It's like a word that has lost most of its sharp edge.
  • The "N-Word": This word carries a heavy history of violence and trauma. The disagreement here was huge. The AI was very aggressive in flagging it, often hurting Black users who were using it among themselves. The humans were much more nuanced, understanding that context is everything.

The Takeaway: We Need a "Human-in-the-Loop"

The paper concludes that we cannot rely on a simple "Gold Standard" (a single set of rules) for moderation.

The Metaphor:
Trying to moderate reclaimed language with a simple AI is like trying to judge a complex family argument by only looking at the volume of the voices. You miss the tone, the history, and the relationship.

What should we do?

  1. Stop assuming everyone in a community thinks alike. We need to accept that people will disagree, and that disagreement is normal.
  2. Context is King. We need systems that can read the room, not just the words.
  3. Don't let the robot be the only judge. We need human nuance, especially from the communities being affected, to decide what is actually harmful.

In short: The paper says, "Hey, AI, you're doing a bad job at understanding the complex, messy, beautiful, and painful ways humans use language to connect. You're hurting the people you're trying to help. We need to do better."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →