On Defining Erasure Harms for NLP
This paper addresses the lack of a cohesive conceptual foundation for identifying and measuring representational harms in NLP by proposing a structured definition of "erasure" that clarifies the necessary components for practitioners to explicitly articulate and operationalize the phenomenon across diverse settings.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are looking at a photograph of a busy street. If someone takes a pair of scissors and cuts out the faces of a specific group of people from that photo, you have "erasure." You can see the street, but a whole group of people is missing, or their presence is so faint you barely notice them.
This paper is about how we, as builders and users of AI language systems (like chatbots or news summarizers), need to stop just guessing when this "cutting out" happens and start having a clear, shared rulebook for it.
Here is the simple breakdown of what the authors are saying:
The Problem: We Are All Using Different Rulers
Right now, when researchers talk about "erasure" in AI, they are all using different definitions.
- The "Too Broad" Ruler: Some say, "If a group isn't mentioned enough, that's erasure." This is too vague. It's like saying "the room is messy" without defining what "messy" actually looks like. You can't measure it.
- The "Too Specific" Ruler: Others say, "Erasure is when an AI fails to tag images of Indigenous people correctly." This is great for that specific job, but it doesn't help us understand erasure in a news summary or a chatbot.
Because we don't have a standard definition, it's hard to build tools to measure if an AI is actually hurting people by "erasing" them.
The Solution: A Six-Part Recipe
The authors propose a structured "recipe" to define erasure. They say that before you can say an AI has erased something, you must explicitly fill in six specific blanks. Think of this like setting up a court case; you need all the evidence in place before you can make a verdict.
Here are the six ingredients you need to define erasure:
- The Target (Who is missing?): Who or what are we worried about losing? Is it a specific person's identity? A whole culture? A history of a specific event?
- The Object (Where are we looking?): What are we examining? Is it the AI's written summary? A list of search results? A generated story?
- The Background Conditions (Why does it matter?): This is the most important part. You need to identify the "unfair patterns" we are trying to avoid.
- Analogy: Imagine a news report that says, "A shooting happened," instead of "Police shot a man." The "background condition" is the historical pattern of media hiding the role of police in violence. If the AI repeats this pattern, it's bad. If the AI just omits a random fact that doesn't fit a harmful pattern, it might just be a mistake, not erasure.
- The Base Knowledge (What should be there?): What information was supposed to be in the output? If you ask an AI to summarize a text, the "base knowledge" is the original text. If the AI leaves out something that wasn't in the original text, that's not erasure; that's just a summary.
- The Alternative (What would a "good" version look like?): You need to imagine a better version of the output. If the AI wrote a summary that hid the police's role, the "alternative" is a summary that clearly states the police fired the shots.
- The De-emphasis (How was it hidden?): How did the AI make the target less visible compared to the "good" version? Did it use passive voice? Did it leave out a name? Did it put the information at the very bottom of a list?
A Real-World Example from the Paper
The authors use a news summarizer to show how this works.
- The Situation: An article says, "Police officers shot and killed a man."
- The AI Output: The AI summarizes it as, "He died in an officer-involved shooting."
- Applying the Recipe:
- Target: The role of the police.
- Background Condition: We know that media often tries to minimize police violence (a harmful pattern).
- Base Knowledge: The original article clearly stated the police did the shooting.
- Alternative: A summary that says, "Police officers shot and killed him."
- De-emphasis: The AI used passive voice ("officer-involved") to hide who pulled the trigger.
- Verdict: This is erasure. The AI took a clear fact and used language to make the police's role less visible, reproducing a harmful historical pattern.
Why This Matters
The paper argues that you cannot just say "AI is biased" and move on. You have to be specific. By filling in these six blanks, practitioners can:
- Agree on what is happening: Stop arguing about whether something is erasure and start measuring how much it is happening.
- Build better tools: Once you know exactly what you are looking for (e.g., "passive voice hiding police roles"), you can build specific tests to catch it.
- Avoid false alarms: You won't accidentally call something "erasure" just because a fact was missing, if that fact wasn't part of a harmful pattern to begin with.
What the Paper Does NOT Say
- It does not claim that fixing erasure is easy.
- It does not say that AI is the only cause of these problems (humans write the news, humans train the AI).
- It does not offer a magic fix or a specific software tool to solve this today. It only offers a definition and a framework to help people think about the problem more clearly.
In short, this paper is a call to stop using the word "erasure" loosely. It asks us to put on our detective hats, identify the six clues, and only then declare that an AI has truly erased someone or something.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.