Labels or Input? Rethinking Augmentation in Multimodal Hate Detection
This paper proposes an end-to-end pipeline that enhances compact Vision-Language Models for multimodal hate detection by combining structured prompts, granular labeling, and a novel counterfactual data augmentation framework, thereby achieving robust performance on nuanced content without relying on costly large models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet is a giant, chaotic digital town square. In this square, people share jokes, news, and memes. But sometimes, hidden inside a funny picture with a silly caption, there's a mean-spirited insult or a hateful message. This is like a "wolf in sheep's clothing."
For a long time, computers have tried to spot these "hateful memes" to keep the town square safe. But it's a tough job because the hate isn't always obvious; it's often a mix of a harmless picture and a mean joke.
This paper is like a team of detectives trying to figure out the best way to teach computers to spot these hidden wolves without hiring a million expensive experts to do the work. They tested two main strategies: changing the instructions given to the computer and creating better practice tests.
The Two Detective Strategies
1. The "Instruction Manual" Upgrade (Prompt Optimization)
Imagine you are teaching a robot to spot bad jokes.
- The Old Way: You just say, "Is this hateful? Yes or No."
- The New Way: The authors tried giving the robot a much more detailed instruction manual. Instead of a simple "Yes/No," they asked the robot to rate the hatefulness on a scale from 0 to 9 (like a pain scale) and gave it specific categories to look for, like "is this sexist?" or "is this racist?"
What they found:
- If the task is simple (just Yes/No), a short, simple instruction works best.
- But if the task is tricky (rating how bad something is), the detailed, structured instructions made the robot much smarter.
- The Surprise: For the biggest, smartest robots, giving them these detailed instructions worked almost as well as spending months training them on new data. It's like realizing that a smart student doesn't need to re-learn the whole subject; they just need a better study guide.
2. The "Counterfeit" Practice Test (Multimodal Augmentation)
The second problem the team faced was that the "textbooks" (datasets) the robots were studying were flawed. Many hateful memes in the training data were just bad pictures with mean captions stuck on top. The robots learned to associate "mean words" with "hate," even if the picture was innocent.
To fix this, the team built a factory that creates "counterfeit" practice tests.
- How it works: They took a hateful meme (a nice picture + a mean caption) and used AI to swap the mean caption for a nice, neutral one, while keeping the picture exactly the same.
- The Goal: This teaches the robot: "Hey, this picture is actually fine! It's only hateful because of the words. If you change the words, the hate disappears."
What they found:
- By adding these "neutralized" memes to the training data, the robots became much better at understanding the context. They stopped just looking for bad words and started looking at the whole picture.
- This worked especially well for the smarter, larger robots, helping them catch the subtle, tricky jokes that usually slip through the cracks.
The Results: A Better, Cheaper Solution
The paper concludes that you don't always need the most expensive, massive super-computers to solve this problem.
- Better Instructions Matter: If you give a smaller, cheaper computer a really good, structured set of instructions (prompts), it can perform almost as well as a giant, expensive one.
- Better Data Matters: If you teach computers with "cleaned up" examples (where the hate is removed but the picture stays), they learn to be fairer and less likely to make mistakes based on stereotypes.
The Bottom Line
The authors built a system that creates these "cleaned up" memes and better instructions to help smaller, more affordable computers do a great job at spotting online hate. They proved that by fixing how we talk to the AI and what we teach it, we can build safer online spaces without needing to spend a fortune on massive technology.
Important Note: The team was very careful to say they did not create new hateful content. Their "factory" only took existing bad memes and turned them into good, neutral ones to use as teaching tools. Their goal was to reduce harm, not create more of it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.