Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
This paper introduces a novel multimodal and multilingual benchmark comprising 3,000 texts, 6,000 images, and 1,200 videos in English and Arabic to evaluate the ability of state-of-the-art models to distinguish between safe, overtly harmful, and covertly harmful humor, revealing significant performance gaps between closed- and open-source models and highlighting the urgent need for culturally grounded safety alignment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to understand human jokes. You show it a picture of a cat wearing a hat, and it laughs. Easy! But then you show it a joke that sounds funny on the surface but is actually mean-spirited, or a meme that relies on a specific cultural inside joke. Suddenly, the robot gets confused. It might think a harmless joke is dangerous, or worse, it might laugh at a joke that is actually hurtful.
This paper, "Harm or Humor," is like a giant, tricky "final exam" designed to test exactly how good AI is at spotting the difference between a funny joke and a harmful one.
Here is the breakdown of the paper using some everyday analogies:
1. The Problem: The Robot's "Cultural Blindness"
Think of AI models as newly hired interns who have read every book in the library but have never actually lived in the real world. They know the dictionary definitions of words, but they don't understand nuance.
- The "Overt" Harm: This is like a clown wearing a sign that says "I am being mean." The AI can easily spot this because the "mean" part is written in big, bold letters.
- The "Covert" (Implicit) Harm: This is like a clown whispering a joke that only people from a specific village understand. To an outsider, it sounds innocent. To the villagers, it's a cruel insult. Current AI models often miss these whispers because they lack cultural context and the ability to "read between the lines."
2. The Solution: The "Harm or Humor" Exam
The researchers created a massive, three-part test to see if AI can pass the "Cultural Intelligence" course. They didn't just use text; they used a mix of media, just like how we consume humor in real life.
- The Text Section: 3,000 jokes (in English and Arabic).
- The Image Section: 6,000 memes (pictures with text).
- The Video Section: 1,200 short clips (where timing, voice, and visuals all matter).
The Twist: They didn't just ask, "Is this funny?" They asked, "Is this safe?" And if it's not safe, is the danger obvious (like a slapstick fall) or hidden (like a subtle stereotype)?
3. The Test Subjects: The "Students"
They tested two types of AI "students":
- The Private Schools (Closed-Source Models): These are the big, expensive AI models from companies like OpenAI (GPT) and Google (Gemini). They have seen a lot of data and have strict rules.
- The Public Schools (Open-Source Models): These are models anyone can download and study. They are often more flexible but sometimes less disciplined.
4. The Results: Who Passed?
The results were a bit of a shock, like a student acing a math test but failing a test on local customs.
- The "English" Bias: The AI models were much better at understanding jokes in English than in Arabic. It's like a student who is fluent in French but stumbles over simple Spanish idioms. Even the smartest models struggled with the cultural nuances of Arabic humor.
- The "Implicit" Trap: The models were great at spotting the "clowns with signs" (Explicit harm). But when it came to the "whispered jokes" (Implicit harm), they often failed. They missed the hidden insults because they were looking for surface-level keywords rather than deep meaning.
- The "Safety Overreaction": Some open-source models were so scared of making a mistake that they refused to answer or labeled everything as "safe." It's like a security guard who, afraid of letting a bad guy in, locks the door and keeps everyone out. This is bad because it means they miss the actual harmful jokes.
5. The Big Takeaway
The paper concludes that making AI bigger isn't enough. You can't just feed an AI more data and expect it to understand human culture.
To make AI safe, we need to teach it cultural empathy. We need it to understand why a joke hurts, not just what words are used. Currently, AI is like a tourist who knows the map but doesn't know the local customs; it needs to learn how to be a local to truly understand humor and safety.
In short: The paper built a tough test to show that while AI is getting smarter, it still struggles to understand the "unspoken rules" of humor, especially in languages and cultures other than English. Until we fix this, AI might keep accidentally laughing at things it shouldn't, or getting offended by things it shouldn't.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.