Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed
This position paper argues that universal toxicity detectors fail to protect marginalized communities, such as those with dwarfism or visual impairments, and demonstrates that community-specific toxicity detection, while significantly improving harm recognition through prompt-based adaptation and fine-tuning, still requires substantial research to reach the performance levels of general-purpose systems.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are walking through a giant, magical art gallery where the walls are made of living paint. This isn't just any gallery; it's powered by a super-smart robot artist that can turn your wildest words into pictures instantly. If you say, "a cat wearing a hat," it paints a cat in a hat. This technology, called Text-to-Image generation, is like a digital genie that grants visual wishes. But, just like any powerful tool, it can sometimes make mistakes. Sometimes, it might paint something scary, violent, or just plain rude. To stop this, scientists built "safety guards"—digital bouncers that look at every picture before it goes on the wall. If the bouncer sees something bad, it blocks the picture.
For a long time, these bouncers have been trained with a "one-size-fits-all" rulebook. They are taught to spot obvious bad stuff like violence or hate speech, which is like teaching a security guard to only look for people carrying weapons. But what if the danger isn't a weapon? What if the danger is a subtle, hurtful stereotype that only a specific group of people would recognize? This paper dives into that exact problem. It asks: What if the safety guard is so busy looking for swords that it misses the fact that it's painting a wheelchair user with wheels for legs, or a blind person with a blindfold over their eyes while they are reading a book? The researchers argue that we need a new kind of safety guard—one that understands the specific, unique rules of different communities, rather than just applying the same generic rules to everyone.
The Problem: The "Universal" Bouncer is Asleep at the Wheel
The researchers started by testing the current top-tier safety guards (the "State-of-the-Art" detectors) on images generated for two specific communities: people who are blind or have low vision (BLV), and people with dwarfism (DWF). They created a set of 2,400 images based on real prompts from these communities, asking the AI to show them doing everyday things like cooking, working, or playing.
Then, they asked five disability experts to look at these images and point out anything that felt wrong or harmful. The experts found a shocking gap. Even though the current safety bouncers gave these images a "Safe" stamp of approval, the experts found that 35% of them were actually harmful.
To give you a concrete example of what "harmful" looks like here:
- For the Blind/Low Vision community: The AI might draw a person with robotic, metallic eyes, or show a guide dog wearing a blindfold (which is silly and confusing), or depict a walking stick being used like a cooking spoon.
- For the Dwarfism community: The AI might draw an adult with dwarfism as a baby, turn them into a fantasy character like a dwarf from a fairy tale, or show women with dwarfism wearing beards.
These aren't just "oops" moments; they are deep, stereotypical errors that reinforce harmful ideas. The paper shows that the current "universal" detectors are completely blind to these specific types of harm. When the researchers asked these detectors to look at the images using the new, specific rules, the detectors failed miserably. In fact, their performance was often worse than random guessing. It's like asking a bouncer who only knows how to spot guns to find a fake mustache; they just don't have the right tools for the job.
The Solution: Teaching the Guard the Specific Rules
Since the "universal" bouncer wasn't working, the team tried to teach the AI new tricks without rebuilding the whole robot from scratch. They tested three different ways to adapt the safety guards to these specific communities:
- The "Show and Tell" Method (In-Context Learning): Imagine showing the AI a few examples of bad pictures and saying, "See this? This is bad because of X. Now look at this new picture." This method helped a little bit, improving the detection scores, but it was inconsistent. Sometimes the AI got it right, and sometimes it still got confused.
- The "Interrogation" Method (Visual Question Answering): Instead of just asking "Is this safe?", the researchers asked the AI a series of specific questions like, "Does this person have a beard?" or "Are the eyes realistic?" If the AI answered "Yes" to a bad thing, the picture was flagged. This worked much better! It forced the AI to look closely at the details. For the Dwarfism community, this method boosted the detection score to a much more promising level.
- The "Study Session" (Fine-Tuning): This is where the AI gets a small crash course using a tiny dataset of about 100 examples. They taught the AI the specific rules for these communities. This was the most effective method for smaller AI models, bringing their performance up significantly.
The Catch: It's Harder Than It Looks
While these methods improved things, the paper makes it very clear that we haven't solved the problem yet. Even with the best methods, the AI's ability to spot these specific harms is still far below its ability to spot general bad stuff. The researchers found that the AI is very sensitive to changes. If the community says, "Actually, we don't think that specific thing is harmful anymore," the AI trained on the old rules gets confused and fails to adapt quickly.
The paper concludes that while we have made progress, building a safety system that truly respects and protects marginalized communities is a massive challenge. It's not just about writing a better rulebook; it requires a constant, evolving partnership between the AI developers and the communities themselves. The "one-size-fits-all" approach is officially out; the future of AI safety needs to be a patchwork of specific, community-led understanding. The researchers suggest that we need to keep working on this, because until we can reliably detect these subtle harms, the "safe" images we see might still be reinforcing stereotypes that hurt real people.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.