← Latest papers
💬 NLP

Epistemic Injustice in Language Models: An Audit of Pretraining Filters and Guardrails

This paper audits pretraining filters and inference-time guardrails in language models, revealing that they disproportionately erase mentions of marginalized groups through reliance on simplistic lexical cues while failing to catch actual hate speech or private information, thereby creating a form of epistemic injustice that contradicts human judgment.

Original authors: Marco Antonio Stranisci, A Pranav, Rossana Damiano, Christian Hardmeier, Anne Lauscher

Published 2026-06-05
📖 5 min read🧠 Deep dive

Original authors: Marco Antonio Stranisci, A Pranav, Rossana Damiano, Christian Hardmeier, Anne Lauscher

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are building a massive library of books to teach a super-smart robot how to speak and understand the world. This robot, called a Large Language Model (LLM), needs to read billions of sentences from the internet to learn.

However, before the robot starts reading, humans put up filters (like a strict librarian) to throw out "bad" books. Then, when the robot starts talking to people, humans put up guardrails (like a safety inspector) to stop it from saying anything "unsafe."

This paper is an audit—a deep investigation—into how these librarians and safety inspectors are doing their jobs. The researchers found that while they are trying to keep the robot safe, they are accidentally erasing the voices of marginalized groups (like transgender people, women, and people from Central America) and replacing human judgment with rigid, often mistaken, rules.

Here is a breakdown of their findings using simple analogies:

1. The "Banned Word" Trap

The researchers found that many of these safety systems work like a child's "naughty list."

  • How it works: If a sentence contains a specific "bad" word (like a swear word or a sexual term), the system immediately deletes it, without looking at the context.
  • The Problem: It's like a librarian throwing out a book about a medical procedure just because it mentions a body part, or deleting a story about a sex worker's rights because it uses the word "sex."
  • The Result: The systems are very good at catching words on their "naughty list," but they are terrible at spotting actual harm, like privacy leaks, copyright theft, or real hate speech that doesn't use the specific banned words.

2. The "Over-Protective" Filter

The study discovered that these filters are over-protective against specific groups, acting like a bouncer who is suspicious of everyone wearing a certain type of hat.

  • Transgender People: Sentences mentioning transgender identities were flagged for removal much more often than sentences about cisgender (non-trans) people.
  • Women: Mentions of women were flagged significantly more often than mentions of men.
  • Central Americans: Content mentioning people from Central America was almost always flagged for removal, while content about people from Western Europe was rarely touched.
  • The Analogy: Imagine a security guard at a museum who lets everyone in except for people from one specific neighborhood. The guard isn't trying to be mean; they are just following a flawed rule that assumes people from that neighborhood are "risky." This results in the museum (the AI) never learning about that neighborhood's culture.

3. The Robot vs. The Human

The researchers asked human volunteers to look at the sentences the robots flagged and decide: "Should this stay or go?"

  • The Shocking Stat: The humans said "Keep it" 88.5% of the time for sentences the pre-training filters wanted to delete, and 91.3% of the time for sentences the guardrails wanted to block.
  • The Analogy: It's like a robot chef tasting a soup, deciding it's "poisonous" because it contains a specific herb, and throwing the whole pot away. A human chef tastes it, realizes it's just a spicy soup, and says, "No, this is delicious and safe to eat."
  • The Conflict: The automated systems are removing content that humans actually find valuable, educational, or harmless. By deleting these sentences, the AI is losing the ability to understand these groups and their experiences.

4. The "Silent Erasure"

The paper calls this "Epistemic Erasure."

  • What it means: "Epistemic" relates to knowledge. "Erasure" means wiping something out.
  • The Metaphor: Imagine a map of the world where the AI is drawing the continents. Because the filters keep deleting sentences about transgender people or Central Americans, the AI's map ends up with blank white spots where those groups should be. The AI literally doesn't "know" they exist or what their lives are like because the data was scrubbed away before it could learn.

Summary of the Paper's Claims

  • The Systems Don't Agree: The different filters and guardrails often disagree with each other. Some catch things others miss, and they rely heavily on simple word lists rather than understanding context.
  • Bias is Systemic: This isn't just a glitch; it's a pattern. Systems from different companies and different continents all seem to over-flag the same marginalized groups.
  • Human Judgment is Different: Humans generally see these sentences as safe or important, while the machines see them as dangerous.
  • The Consequence: By relying on these automated systems, we are actively shaping AI to be blind to the realities of transgender people, women, and people from specific regions, effectively silencing them in the digital world.

The paper concludes that to fix this, we can't just rely on automated "naughty lists." We need to involve the people being affected in the decision-making process and create systems that understand context rather than just counting bad words.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →