When Hate Meets Facts: LLMs-in-the-Loop for Check-worthiness Detection in Hate Speech
This paper introduces WSF-ARG+, the first dataset combining hate speech with check-worthiness information, and proposes an LLM-in-the-loop framework that reduces human annotation effort while demonstrating that incorporating check-worthiness labels significantly improves hate speech detection performance.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet as a giant, chaotic town square. In this square, there are two main problems: people shouting hateful slurs (Hate Speech) and people spreading lies or half-truths (Misinformation).
For a long time, researchers treated these as two separate problems. They built tools to catch the shouters and separate tools to catch the liars. But this paper argues that in the real world, these two problems often wear the same mask. Hate speech is increasingly being disguised as "facts" or "news."
Here is a simple breakdown of what the authors did, using some everyday analogies.
1. The Problem: The "Wolf in Sheep's Clothing"
Imagine a person trying to start a riot. Instead of just screaming insults (which is easy to spot), they stand on a soapbox and say, "Did you know that 90% of [Group X] are actually aliens?"
This is a fact-like claim. It sounds like a news report, but it's a lie designed to make people hate a specific group.
- The old way: Moderators might miss this because it doesn't look like a standard insult.
- The new way: If we can spot that this sentence is a "claim" that needs fact-checking, we can catch the hate speech hiding inside it.
The authors call this intersection "Check-worthiness." It asks: "Is this statement a claim that matters enough to be fact-checked?"
2. The Solution: The "Human-LLM Team"
To study this, the researchers needed a massive dataset of these tricky messages. But labeling them is hard work. It's like trying to grade 1,000 essays where the answers are subjective. You need experts, and experts are expensive and tired.
So, they invented a new workflow called "LLM-in-the-Loop."
Think of this like a construction site:
- The LLM (Large Language Model) is the Power Drill. It's fast, strong, and can do the heavy lifting. It scans thousands of messages and says, "This looks like a claim!" or "This is just an opinion!"
- The Human is the Foreman. They don't do every single screw. Instead, they watch the Power Drill.
- If the Drill and the Foreman agree, they move on.
- If the Drill gets confused or disagrees with the Foreman, the Foreman steps in to make the final call.
The Result: They used 12 different "Power Drills" (AI models) to help them build a new dataset called WSF-ARG+. This dataset is special because it contains both hateful and non-hateful messages, and every single "fact" inside them has been labeled as either "Check-worthy" or "Not Check-worthy."
3. The Discovery: Lies Make Hate Worse
Once they had their new dataset, they ran some experiments and found two surprising things:
A. The "Poisoned Chalice" Effect
They found that hate speech containing "check-worthy" claims (lies presented as facts) is more harmful than simple insults.
- Analogy: A simple insult is like a paper cut; it stings but heals fast. A lie wrapped in hate is like a poisoned dart; it spreads fear and causes deeper, longer-lasting damage to the community. The data showed these messages were rated significantly higher for "harassment" and "hate."
B. The "Super-Helper" Effect
They tested if giving AI models a "cheat sheet" (the check-worthiness labels) helped them detect hate speech better.
- Analogy: Imagine trying to find a needle in a haystack. If you tell the AI, "The needle is wrapped in red tape," it finds it much faster.
- When they gave the AI the "check-worthiness" labels, the AI's ability to spot hate speech improved significantly, especially for the larger, smarter models. It went from being a decent detective to a master detective.
4. Why This Matters
This paper is a blueprint for the future of internet safety.
- We have a new map: They released the WSF-ARG+ dataset, which is the first map of where hate speech and fake facts overlap.
- We have a better tool: They proved you don't need to hire 100 humans to label data. You can use a smart AI as a "junior assistant" and have a human "manager" fix the mistakes. This saves time and money without losing quality.
- We understand the enemy better: We now know that when hate speech tries to sound like "science" or "news," it becomes more dangerous. By focusing on the facts inside the hate, we can dismantle the argument more effectively than just blocking the insult.
In short: The authors built a smarter way to catch the liars who are also the haters, using a team of humans and AI working together, and they proved that catching these specific types of lies makes the internet a safer place.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.