SAFE-MEME: Structured Reasoning Framework for Robust Hate Speech Detection in Memes
The paper introduces SAFE-MEME, a structured reasoning framework with Q&A and hierarchical variants, alongside two new fine-grained multimodal hate speech datasets (MHS and MHS-Con), to achieve robust hate speech detection in memes that rivals closed-source models and outperforms existing open-source baselines.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet as a massive, chaotic digital town square where everyone is shouting, whispering, and drawing pictures on the walls. In this town, a "meme" is like a secret handshake made of a picture and a caption. Sometimes, it's just a funny joke about a cat. But other times, it's a Trojan horse: a picture that looks innocent, paired with a caption that seems harmless, but together they carry a hidden, poisonous message meant to hurt a specific group of people. This is called "hate speech," and when it hides inside a meme, it's incredibly hard to catch because the hate isn't in the words alone, or the picture alone—it's in the tricky combination of both.
For a long time, computers trying to police this town square have been like security guards who only read the text or only look at the picture. They often miss the joke that isn't a joke, or the insult that is disguised as a compliment. They struggle with "implicit" hate, which is like a whisper that only makes sense if you know the secret history of the town. The big question researchers have been asking is: Can we teach computers to stop just looking at the surface and start thinking about what the meme really means, just like a human would?
This paper introduces a new team of digital detectives called SAFE-MEME. Instead of just scanning a meme and shouting "Hate!" or "Safe!", SAFE-MEME acts like a curious teenager who refuses to take things at face value. It uses a special "structured reasoning" trick, similar to how a detective solves a mystery by asking a series of questions before drawing a conclusion.
The researchers built two new training grounds for their detectives. The first, called MHS, is a library of memes labeled as "Explicit Hate" (obvious insults), "Implicit Hate" (sneaky, hidden insults), or "Benign" (harmless). The second, MHS-Con, is a stress test designed to trick the computers. It takes the same picture and pairs it with three different captions: one that is clearly hateful, one that is sneakily hateful, and one that is totally innocent. This setup is like showing a detective the same photo of a person and asking, "Is this person a hero, a villain, or just a neighbor?" depending on the story you tell about them.
The paper tests two versions of the SAFE-MEME detective:
- SAFE-MEME-QA: This version plays a game of "Question and Answer." When it sees a meme, it doesn't just guess. It asks itself questions like, "Who is this targeting?" and "What kind of hate is this?" It writes down the answers, building a chain of logic before deciding if the meme is bad.
- SAFE-MEME-H: This version is more like a librarian who first describes the picture in great detail and then sorts the meme into a hierarchy of categories (Is it hateful? If yes, is it obvious or hidden?).
The results are a bit like a surprise party. When the detectives faced the standard library of memes (MHS), the librarian-style detective (SAFE-MEME-H) was the star player. It beat the best open-source computer models by a solid margin, proving that taking a step-by-step, hierarchical approach works well for spotting subtle, hidden hate. However, when they moved to the tricky stress test (MHS-Con), where the same image had to be judged against three very different stories, the Question-and-Answer detective (SAFE-MEME-QA) took the lead. It outperformed the other open-source models by nearly 5% and handled the confusing, "confounding" cases much better than its sibling.
Interestingly, the paper found that the most powerful "closed-source" models (the super-smart, expensive AI systems from big tech companies) were often too cautious. They missed a lot of the sneaky, implicit hate because they were afraid of making a mistake. On the other hand, a simple text-only model (T5-large) actually did surprisingly well on the stress test, but only because the stress test relied heavily on obvious words. The SAFE-MEME detectives, however, showed they could actually understand the context, not just the words.
The researchers also looked at where their detectives made mistakes. Sometimes, the AI got confused by rare groups of people it hadn't seen enough of in its training, or it mixed up different cultural references. It's like a detective who knows a lot about one neighborhood but gets lost in another. The paper concludes that while SAFE-MEME is a huge step forward in teaching computers to "think" about memes rather than just "see" them, it still carries some of the biases from the data it was trained on. It's not a perfect solution yet, but it's a much smarter way to look at the digital town square than we had before.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.