What Are They Filtering Out? An Experimental Benchmark of Filtering Strategies for Harm Reduction in Pretraining Datasets
This paper presents a benchmark study revealing that while data filtering strategies effectively reduce harmful content in pretraining datasets for Large Language Models, they inadvertently increase the underrepresentation of vulnerable groups, thereby exacerbating discrimination.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to build the world's smartest librarian. This librarian (an AI) needs to read millions of books, websites, and articles to learn how to speak and think like a human. But, the internet is a messy place. It's full of great stories, but also full of hate speech, violence, and terrible stereotypes.
To make sure this librarian doesn't learn bad habits, the builders try to "filter" the books before giving them to the librarian. They want to throw away the "bad" books.
This paper is like a detective report investigating what happens when we try to clean up the library. The researchers asked two big questions:
- How are people currently trying to clean the library?
- Who gets accidentally kicked out of the library while we are trying to remove the bad stuff?
Here is the breakdown of their findings, using some simple analogies:
1. The "Cleaning Crew" (The Strategies)
The researchers looked at 55 different manuals from companies building these AI librarians. They found that everyone uses different cleaning tools, but they rarely tell the public exactly how they use them.
Think of the cleaning tools like different types of brooms:
- The "Bad Word" Broom (Rule-based): This broom sweeps up anything with specific "dirty" words on it (like swear words or slurs).
- The "Smell Test" Broom (Toxicity Classifiers): This is a robot nose that sniffs a page. If it smells "toxic," it throws the page away.
- The "Quality" Broom (Document Similarity): This broom compares a page to a "perfect" page (like a Wikipedia article). If the new page doesn't look similar enough to the perfect one, it gets tossed.
- The "Blacklist" Broom: This just blocks entire websites known for being trouble.
The Problem: The researchers found that most companies are being very secretive. They say, "We cleaned it!" but they won't show their work. It's like a chef saying, "This soup is safe to eat," but refusing to show you the recipe or the ingredients.
2. The Big Surprise: The "Collateral Damage"
The most important part of the paper is the experiment. The researchers took a huge pile of internet text and ran it through these different "brooms" to see what got thrown away.
They discovered a scary side effect: In trying to remove the "bad" stuff, the cleaning crew is accidentally throwing away the stories of vulnerable people, especially women.
Here is the metaphor:
Imagine you are cleaning a garden to remove weeds (harmful content).
- The Goal: Remove the poisonous weeds.
- The Mistake: The "Quality" broom is so aggressive that it thinks any flower that looks a little different from the "perfect" rose is a weed.
- The Result: You end up with a garden full of strong, standard-looking plants (often representing men or Western perspectives), but you have accidentally pulled out almost all the rare, beautiful flowers that represent women and people from post-colonial backgrounds.
3. What Did They Find Specifically?
- Women get filtered out the most: When the AI uses "Bad Word" filters or "Smell Test" robots, mentions of women disappear from the data much faster than mentions of men.
- Different brooms catch different things:
- One broom catches pornography words.
- Another catches racist slurs.
- Another catches "low quality" writing.
- The lesson: If you pick one broom, you only catch one type of bad thing, but you might miss other types of bad things, or you might accidentally throw away good things that look like the bad things.
- The "Quality" Trap: The researchers found that the "Quality" broom (which compares text to Wikipedia) is actually terrible at finding hate speech. It keeps a lot of toxic content but throws away a lot of unique voices. It's like judging a book by its cover and throwing away a brilliant story just because the font looks weird.
4. The "Occupation" Clue
The researchers looked at what was being thrown out.
- When they filtered for "bad" content, they found that mentions of women working as actors or pornographic actors were removed at a very high rate.
- Mentions of men working as politicians or writers were kept much more often.
- This means the AI is learning a distorted view of the world where women are mostly seen in sexualized or removed contexts, while men are seen as leaders and creators.
The Bottom Line
The paper concludes that we are trying to fix the AI, but we are breaking the balance.
By trying to make AI "safe" by filtering out harmful content, we are inadvertently making the AI less representative of real life. We are creating a digital world where women and minorities are invisible because the "safety filters" are too blunt.
The Takeaway: We can't just use a sledgehammer to clean up the internet. We need a much more delicate, surgical approach that knows the difference between a "weed" (hate speech) and a "rare flower" (a unique perspective from a vulnerable group). If we don't fix this, our AI librarians will only know how to talk about half the world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.