← Latest papers
💬 NLP

Large Language Models in the Abuse Detection Pipeline

This survey paper analyzes the integration of Large Language Models into the four stages of the Abuse Detection Lifecycle—Label & Feature Generation, Detection, Review & Appeals, and Auditing & Governance—while synthesizing current practices, architectural considerations, and key challenges like latency and fairness to guide the deployment of reliable, large-scale safety systems.

Original authors: Suraj Kath, Sanket Badhe, Preet Shah, Ashwin Sampathkumar, Shivani Gupta

Published 2026-04-03
📖 5 min read🧠 Deep dive

Original authors: Suraj Kath, Sanket Badhe, Preet Shah, Ashwin Sampathkumar, Shivani Gupta

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet as a massive, bustling city. In this city, there are millions of people chatting, sharing photos, and posting news every second. But like any big city, there are also troublemakers: people spreading lies, bullying others, running scams, or trying to trick the system.

For a long time, the "City Watch" (the platforms like Google, Facebook, or X) used a very simple way to catch these troublemakers. They had a list of "bad words" (like specific insults). If a post had a bad word, the computer flagged it. If it didn't, it was safe.

The Problem: The troublemakers got smarter. They stopped using obvious bad words. Instead, they used sarcasm, inside jokes, coded language, or subtle threats. The old "bad word" list couldn't catch them. Also, the city grew so big that human guards couldn't read every single post.

The New Solution: Enter the Large Language Models (LLMs). Think of these not as simple spell-checkers, but as super-intelligent, well-read detectives who understand context, culture, and nuance. They can read a post, understand the tone, and say, "This isn't just a joke; this is actually a threat."

This paper is a guidebook on how to use these super-detectives to run the entire "City Watch" operation. The authors break the job down into four main stages, like a factory assembly line for safety.

The Four Stages of the Safety Factory

1. The Training Room (Label & Feature Generation)

The Challenge: Before a detective can catch a criminal, they need to study past cases. But in the internet city, new types of crimes appear every day (like new slang for bullying). Humans are too slow to write down examples of these new crimes.
The LLM Solution: The LLM acts as a creative writing coach. It can invent thousands of fake examples of new types of bullying or scams to train the other guards. It can also help humans label old posts faster.

  • The Catch: Sometimes the AI coach gets it wrong or is biased. If the AI is too scared to say anything bad, it might miss real crimes. If it's too aggressive, it might think a harmless joke is a crime. So, humans still need to check the AI's homework.

2. The Front Gate (Detection)

The Challenge: The city gate is flooded with millions of people walking in every second. You can't stop every single person to ask them deep, philosophical questions; the line would never move.
The LLM Solution: You use a two-tier security system.

  • The Bouncers (Small Models): Fast, cheap, and good at spotting obvious bad guys. They scan 99% of the traffic.
  • The Detectives (LLMs): When the Bouncer is confused (e.g., "Is this a joke or a threat?"), they hand the case to the LLM Detective. The Detective takes a moment to think, "Hmm, this person is using sarcasm to insult a specific group. That's a violation."
  • The Catch: The Detective is slow and expensive. You can't use them for everyone. Also, bad guys are now using AI to write their own "jailbreak" scripts to trick the Detective.

3. The Courtroom (Review & Appeals)

The Challenge: When someone is banned, they often say, "But I didn't do anything wrong!" Human judges are overwhelmed and tired. They might make mistakes or give vague reasons like "You violated policy."
The LLM Solution: The LLM acts as a legal assistant. It reads the whole conversation, summarizes the evidence, and writes a clear, polite letter to the banned user explaining exactly why they were caught. It helps the human judge make a fairer, faster decision.

  • The Catch: Sometimes the AI assistant writes a very convincing explanation that sounds logical but is actually wrong (it "hallucinates"). Humans need to make sure the AI isn't just making things up to sound smart.

4. The Inspector General (Auditing & Governance)

The Challenge: Over time, the rules change, and the city's culture shifts. A system that was fair last year might be unfair today. Also, the AI might start treating different groups of people differently (bias).
The LLM Solution: The LLM acts as a stress-tester. It tries to trick the system on purpose. It asks, "Can I say this racist thing without getting caught?" or "Does the system treat men and women differently?" It constantly checks the system for holes and biases to make sure the "City Watch" is fair and honest.

The Big Paradox: The Sword and the Shield

The paper points out a scary twist: The same tool that helps the police is also helping the criminals.

  • The Shield: Defenders use AI to catch bad guys.
  • The Sword: Bad guys use AI to write better scams, create fake news, and write messages that sound perfect to trick the AI guards.

The Future: What Needs to Happen?

The authors say we can't just turn on the AI and hope for the best. We need to fix three big problems:

  1. Speed vs. Smarts: We need to make the AI faster and cheaper so it doesn't slow down the internet.
  2. Honesty: The AI needs to be honest about why it made a decision, not just give a fake reason that sounds good.
  3. Human Control: We can't let the AI run the show alone. Humans must stay in the loop to check the AI's work, fix its biases, and update the rules when the bad guys get smarter.

In short: Large Language Models are like giving our internet safety guards superpowers. They can read between the lines and understand complex situations. But if we aren't careful, those superpowers can also be used to break the system. The key is to use them as powerful assistants, not as the only bosses, and to keep humans in charge to ensure fairness and safety.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →