← Latest papers
💬 NLP

Context-Aware Detection and Victim-Centered Response Generation for Online Harassment in Private Messaging

This study presents a context-aware, victim-centered AI framework that leverages large language models trained on a dataset of private Instagram messages to more accurately detect online harassment and generate psychologically supportive responses that outperform human reactions in terms of emotional support and de-escalation.

Original authors: Pinxian Lu, Nimra Ishfaq, Emma Win, Morgan Rose, Sierra R Strickland, Candice L Biernesser, Jamie Zelazny, Munmun De Choudhury

Published 2026-05-20
📖 5 min read🧠 Deep dive

Original authors: Pinxian Lu, Nimra Ishfaq, Emma Win, Morgan Rose, Sierra R Strickland, Candice L Biernesser, Jamie Zelazny, Munmun De Choudhury

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine your private text messages as a secret diary that only you and your friends can read. Now, imagine someone using that diary to say mean, scary, or hurtful things. This paper is about building a smart helper (an AI) that can read those secret diaries to spot the bullying and then help the person being bullied figure out what to say back.

Here is the story of what the researchers did, broken down into simple parts:

1. The Problem: The "Private Room" vs. The "Public Square"

Most computer programs designed to stop online bullying are like security guards standing in a busy public square (like Twitter or Facebook). They look at loud, public shouts.

But bullying often happens in private rooms (like Instagram Direct Messages). In these private rooms, the bullying is sneaky. It's not just one mean word; it's a whole conversation where the meaning changes based on what was said five minutes ago. It's like a game of "telephone" where the joke turns into a threat only if you know the whole story. Because these conversations are private, computers usually can't see them, and they struggle to understand the context.

2. The Data: A Glimpse into Real Lives

To teach their computer, the researchers asked 26 teenagers (ages 12 to 18) to share their private Instagram messages. Some of these teens were going through very hard times, including thoughts of suicide.

  • The Collection: They gathered over 80,000 messages.
  • The Labeling: Humans read these messages to find the 41 specific instances where bullying actually happened. They had to read the whole conversation to understand if a message was bullying or just a joke between friends.

3. Part One: The "Detective" (Finding the Bullying)

The researchers tried to teach an AI to be a detective.

  • The Old Way: They tried using standard "toxicity detectors" (like a metal detector at an airport). These detectors are good at spotting loud, obvious bad words in public posts, but they failed in private chats because they didn't understand the context.
  • The New Way (The Cascade): They built a two-step "Detective Team."
    1. Detective A reads the message and the 50 messages before it. If it looks suspicious, it flags it.
    2. Detective B is the "Chief Inspector." They look at Detective A's flag and ask, "Are you sure? Is this actually bullying, or just a weird joke?"
  • The Result: This two-step team was much better at finding the bullying than the old detectors. They caught 60% of the bullying cases while keeping false alarms low. The key was that the AI read the whole story, not just the last sentence.

4. Part Two: The "Coach" (Helping the Victim)

Once the AI spots the bullying, the next question is: What should the victim say back?
The researchers built a second AI that acts like a supportive coach.

  • The Strategy: Instead of just writing a random reply, the AI first decides on a strategy (like "be calm," "set a boundary," or "show empathy"). Then, it writes a message based on that strategy.
  • The Style Check: The AI was told to sound like a teenager (using slang, short sentences, and casual text) so it wouldn't sound like a robot or a teacher.
  • The Test: The researchers showed 100 bullying scenarios to human judges. They showed them two options:
    1. What the real teenager actually wrote back.
    2. What the AI wrote back.

The Verdict:

  • Helpfulness: The judges overwhelmingly preferred the AI's responses. They felt the AI's replies were better at stopping the fight, calming things down, and making the victim feel supported.
  • Naturalness: However, the judges felt the real teenagers' replies sounded more "natural" and authentic. The AI was great at being helpful, but sometimes a little stiff compared to how a real teen talks.

5. The Big Idea: "Synthetic Buffering"

The paper introduces a cool concept called "Synthetic Buffering."
Imagine you are in a heated argument and you are so angry or scared that your brain freezes. You don't know what to say.
The AI acts like a cushion or a shield. It takes the emotional weight of the situation, figures out the best thing to say, and drafts a response for you. You can then read it, tweak it, or send it. It doesn't replace you; it helps you find your voice when you've lost it.

Summary of What They Claim

  • Context is King: You can't find bullying in private chats without reading the whole conversation.
  • Two-Step Detection: A team of two AIs works better than one to spot bullying without crying wolf.
  • AI as a Coach: AI can generate replies that are more helpful and calming than what victims usually come up with in the heat of the moment.
  • The Trade-off: AI replies are very helpful but sometimes sound less "real" than human replies.

Important Note: The paper is careful to say this is a tool to help, not a replacement for human friends, parents, or professional help. It's about giving a victim a "just-in-time" boost when they are being hurt online.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →