ToxSyn-PT: A Synthetic Fine-Grained Dataset of Minority-Targeted Toxic Language in Portuguese
This paper introduces ToxSyn-PT, the first large-scale, synthetic Portuguese dataset featuring fine-grained, multi-label hate speech annotations across nine minority groups with essential non-toxic counterexamples, revealing that models trained on social media data fail to generalize to minority-specific contexts and highlighting the limitations of standard evaluation metrics.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to spot a bully in a playground. For a long time, researchers have only been able to teach this robot using examples from English-speaking playgrounds, and even then, they mostly showed the robot the obvious, screaming bullies. They rarely showed the robot the subtle, sneaky bullies who hide their mean words behind jokes or "just asking questions."
Now, imagine trying to teach this robot to understand Portuguese. The problem is, there are almost no good textbooks (datasets) for Portuguese that show the difference between a mean-spirited attack on a specific group of people and a normal, friendly conversation about those same people.
This paper introduces ToxSyn-PT, a brand new, massive "textbook" created specifically to fix this problem. Here is how the authors built it and what they discovered, explained simply:
1. The Problem: The "One-Size-Fits-All" Failure
Think of existing Portuguese hate-speech datasets like a library that only has books about Twitter. If you train your robot using only Twitter books, it becomes an expert at spotting the loud, angry shouting common on social media.
However, the authors found that when they took this "Twitter-trained" robot and put it in a different room (like a news comment section or a formal discussion), it became completely blind. It couldn't recognize the hate speech there because the "language of the bully" changes depending on the room.
- The Analogy: It's like teaching a student to recognize a "wolf" only by looking at wolves in a zoo. When that student goes into the forest, they don't recognize the wolf because it looks different in the wild.
2. The Solution: Building a Synthetic "Training Gym"
Since there weren't enough real examples of Portuguese hate speech targeting specific groups (like Black people, women, LGBTQIA+ individuals, the elderly, etc.), the authors decided to build their own examples using AI.
They created a four-step assembly line (a pipeline) to generate 53,000 sentences:
- Step 1: The Seed. They started with a tiny, high-quality collection of human-written examples (both mean and nice) about two groups.
- Step 2: Expansion. They used an AI (GPT-4) to look at those seeds and write thousands of new variations, targeting nine different minority groups.
- Step 3: Paraphrasing. They took those new sentences and rewrote them in different styles. Crucially, they didn't just write mean things; they wrote nice things about the same groups. This is vital because it teaches the AI the difference between hating a group and talking about a group.
- Step 4: Enrichment. They added even more complex, subtle forms of prejudice (like "ambiguous prejudice" where the hate is hidden in a double-meaning) to make the training harder and better.
The Result: A massive dataset that is perfectly balanced. It has mean sentences, nice sentences, and neutral sentences, all labeled with exactly who is being targeted and what "rhetorical trick" (like sarcasm or victim-blaming) is being used.
3. The Big Discovery: "Catastrophic Failure"
The authors tested their new AI models against old ones. The results were shocking:
- Old Models: When they tried to use models trained on general social media to detect hate speech in this new, specific dataset, they failed miserably. They missed almost all the hate.
- New Models: When they trained models on ToxSyn-PT, those models were great at spotting hate in their own "gym," but they also failed when tested on the old social media data.
The Metaphor: It's like training a swimmer to race in a pool. They become a world-class pool swimmer. But if you put them in the ocean, they drown. Conversely, if you train them in the ocean, they can't swim in the pool. The "water" (the context) is just too different.
This proves that hate speech is not one single thing. It changes based on who is being attacked and where the conversation is happening. The authors also warn that standard "scorecards" (like Macro F1) can be misleading; a robot might get a high score overall but still fail completely at its most important job: spotting the specific hate it was supposed to find.
4. Why This Matters
This paper doesn't claim to have solved hate speech forever. Instead, it provides the first-ever high-quality "gym" for Portuguese AI to learn how to distinguish between genuine hate and normal discussion about minorities.
- Before: Researchers had to guess or use tiny, unbalanced datasets that missed the nuance.
- Now: They have a massive, balanced resource that includes the "good" examples needed to teach AI what not to flag as hate.
The authors released this dataset publicly so other researchers can use it to build better, more careful AI systems that understand the subtle, dangerous ways hate speech targets vulnerable communities in Portuguese.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.