← Latest papers
🤖 machine learning

Detecting LGBTQ+ Instances of Cyberbullying

This study evaluates the effectiveness of various transformer models in accurately detecting and distinguishing cyberbullying incidents specifically targeting the LGBTQ+ community using real social media data.

Original authors: Arslan Bisharat, Manuel Sandoval Madrigal, Mohammed Abuhamad, Deborah L. Hall, Yasin N. Silva

Published 2026-06-10
📖 4 min read☕ Coffee break read

Original authors: Arslan Bisharat, Manuel Sandoval Madrigal, Mohammed Abuhamad, Deborah L. Hall, Yasin N. Silva

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine social media as a giant, bustling town square. In the past, if someone wanted to be mean, they had to stand right next to you to shout insults. Today, the "town square" has a keyboard, and people can throw digital rocks from miles away. This is cyberbullying.

While everyone can get hurt by these digital rocks, the paper explains that a specific group of people—the LGBTQ+ community—is getting hit much harder and more often than others. It's like they are standing in a target zone where the bullies are aiming specifically at their identity.

Here is what the researchers did, explained simply:

The Mission: Building a Digital Security Guard

The researchers wanted to build a "smart security guard" (a computer program) that could look at comments on Instagram and instantly say, "Hey, this is bullying, and it's specifically targeting LGBTQ+ people."

They knew that regular security guards often make mistakes. They might think a comment is just a joke when it's actually hate speech, or they might get confused by slang and hidden insults. So, the team tested three different "brains" (computer models) to see which one was the best guard:

  1. RoBERTa
  2. BERT
  3. GPT-2

The Training Ground: A Small, Tricky Class

To teach these guards, the researchers used a specific set of 1,083 Instagram comments.

  • The Problem: Most of the comments were not about LGBTQ+ bullying. Only about 200 of them were. It's like trying to teach a dog to find a specific rare flower in a field full of dandelions. The dog might just learn to ignore the rare flower because there are so many dandelions.
  • The Fix: To help the guards learn, the researchers used a technique called "oversampling" (SMOTE and ADASYN). Imagine this as taking the few rare flower examples and making photocopies of them so the guards have more practice spotting them.

The Results: Who Was the Best Guard?

After training, they put the guards to the test. Here is what happened:

  • RoBERTa was the champion: It was the smartest guard. It got the highest overall score and was the best at spotting the bullying.
  • The others struggled: BERT and GPT-2 were good at spotting general bullying, but they got confused when the bullying was specifically about LGBTQ+ issues.

The Big Catch: The "Missed Spot" Problem

Even though RoBERTa was the best, it still had a major blind spot.

  • The Good News: It was excellent at saying, "This is not LGBTQ+ bullying." (It rarely cried wolf).
  • The Bad News: It was still bad at saying, "This is LGBTQ+ bullying." It missed a lot of the actual attacks.

Why did it miss them?
The paper explains that bullying against LGBTQ+ people is often like a secret code.

  • Sometimes it uses sarcasm.
  • Sometimes it uses slang that changes meaning depending on the situation.
  • Sometimes it hides the insult inside a sentence that looks innocent.

The computer models are like students who memorized the dictionary but don't understand the joke or the tone. They see the words but miss the sting. For example, the paper found that the model sometimes missed the word "dyke" because it didn't have enough examples of that specific slur in its training, or it missed a comment because it didn't understand the context of the conversation.

The Takeaway

The researchers concluded that while we have built some very smart "security guards" (AI models), they aren't perfect yet.

  • RoBERTa is the strongest guard we have right now.
  • Oversampling (making more practice examples) helped a little, but it didn't fix the problem completely.
  • The Challenge: The "bullying" is too subtle and tricky for the current technology to catch every time. The models are still missing too many of the attacks (False Negatives).

The paper suggests that to make these guards truly effective, we need to teach them with more diverse examples, show them pictures and videos (not just text), and help them understand the deeper context of human conversation. Until then, the digital town square still has blind spots where the LGBTQ+ community remains vulnerable.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →