← Latest papers
💬 NLP

PEACE 2.0: Grounded Explanations and Counter-Speech for Combating Hate Expressions

PEACE 2.0 is a novel tool that enhances hate speech analysis and mitigation by utilizing a Retrieval-Augmented Generation pipeline to provide evidence-grounded explanations and automatically generate factual counter-speech for both explicit and implicit hateful messages.

Original authors: Greta Damo, Stéphane Petiot, Elena Cabrio, Serena Villata

Published 2026-02-20
📖 4 min read☕ Coffee break read

Original authors: Greta Damo, Stéphane Petiot, Elena Cabrio, Serena Villata

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet as a giant, bustling town square. Recently, this square has become overrun with people shouting insults, spreading rumors, and trying to hurt others with their words. This is hate speech.

For a long time, the "security guards" of this town square (the computer programs we use) have been really good at spotting the obvious bullies. If someone yells, "I hate you because of your skin color," the guard immediately points and says, "Stop!"

But here's the problem: The bullies have gotten smarter. They aren't just shouting; they are whispering, using code words, or making jokes that only some people understand. This is implicit hate speech. It's like a bully saying, "Some people just don't belong here," without ever naming who they mean. It's much harder to catch.

Furthermore, even when the guards catch a bully, they usually just throw them out (delete the post). But what if we could teach the bully a lesson, or help the bystanders understand the truth? That's where Counter-Speech comes in—answering hate with facts and kindness instead of just silence.

Enter PEACE 2.0. Think of this tool as a super-powered, high-tech "Peacekeeper" for our town square. It doesn't just catch the bullies; it explains why they are wrong and helps write a perfect, fact-based reply to stop the argument.

Here is how PEACE 2.0 works, broken down into three simple superpowers:

1. The "Fact-Checker" Backpack (RAG)

Imagine you are a student in a debate. If you try to argue without your textbooks, you might make things up or sound unsure. But if you have a backpack full of the world's most trusted books (like the UN's human rights laws), you can win any argument with facts.

PEACE 2.0 has this backpack. It uses a system called RAG (Retrieval-Augmented Generation).

  • Without the backpack: The computer tries to guess a reply based on what it "thinks" is true.
  • With the backpack: The computer instantly searches its library of 30,000+ trusted documents, finds the exact facts needed, and uses them to build its reply.

This means when PEACE 2.0 responds to hate, it's not just guessing; it's citing the law and history.

2. The "Detective's Magnifying Glass" (Explanations)

Sometimes, a message looks innocent but is actually a trap. PEACE 2.0 acts like a detective with a magnifying glass.

  • It reads a message and decides: "Is this hate speech?"
  • Then, it doesn't just say "Yes." It writes a note explaining exactly why.
  • The Magic: Because of the "Fact-Checker Backpack," the detective doesn't just say, "This feels mean." It says, "This sentence is hate speech because it violates Article 12 of the Human Rights Declaration, which protects group X." It connects the dots between the rude comment and the actual rules of the town.

3. The "Diplomat's Pen" (Generating Responses)

This is the newest and coolest feature. Once PEACE 2.0 spots the hate and finds the facts, it writes a response for you.

  • It doesn't write an angry rant back.
  • It writes a Counter-Speech: a calm, persuasive, and fact-filled reply that shuts down the hate without starting a fight.
  • It can show you two versions: one written by a "smart guesser" (no facts) and one written by the "expert with the backpack" (facts). The expert version is always better, more convincing, and harder to argue with.

Why Does This Matter?

The researchers tested PEACE 2.0 against the old way of doing things.

  • The Result: When PEACE 2.0 used its "Fact-Checker Backpack," its explanations were much clearer, and its replies were much more persuasive.
  • The Analogy: It's the difference between a neighbor saying, "That's not true, I think so," versus a neighbor saying, "That's not true. According to the official census data from 2024, the population is actually..." The second one wins the argument every time.

The Bottom Line

PEACE 2.0 is a tool that helps us move from just hiding hate speech to understanding and countering it. It combines the brainpower of a detective, the knowledge of a librarian, and the diplomacy of a peacemaker.

Its goal is to make our online town square a place where we don't just delete the bad words, but where we can actually have better, more informed conversations that make everyone feel safer and included. It's not just about stopping the noise; it's about bringing the peace.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →