Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
This survey provides a comprehensive overview of computational humor in multimodal large language models, organizing the field by capabilities ranging from recognition to generation, synthesizing current datasets and evaluation protocols, and identifying key challenges such as cultural limitations and safety concerns.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are walking through a giant, bustling library where the books aren't just words on a page, but pictures that talk, sing, and tell jokes. This is the world of Multimodal Large Language Models (MLLMs). Think of these AI systems as super-smart students who have read almost every book and looked at almost every picture in the library. They are incredibly good at describing what they see: "That is a dog," or "That is a burning house." They can match a picture of a cat to the word "cat" with perfect accuracy.
But here is the tricky part: human communication isn't just about facts. It's about jokes, sarcasm, and memes. A meme might show a cat sitting in a box, but the joke isn't that the cat is in a box; the joke is that the cat looks like a king ruling a tiny kingdom, or that the box is actually a spaceship. To get the joke, you need to understand the hidden meaning, the cultural references, and the fact that the picture is pretending to be something it isn't. This paper explores why our super-smart AI students are still terrible at getting these jokes, even though they can describe the picture perfectly. It's like having a robot that can list every ingredient in a cake but has no idea why the cake is funny or why it tastes good.
The Great Joke-Understanding Gap
This paper is a massive report card for AI on the subject of multimodal humor. The authors, a team of researchers from universities, looked at how well AI can understand funny images, cartoons, and memes. They found that while AI has gotten really good at the "easy" stuff—like spotting a dog or a car—it is still struggling with the "hard" stuff: figuring out why something is funny.
The researchers organized the problem into three levels, like climbing a ladder:
Level 1: Recognition (The "What")
This is the bottom rung. Can the AI tell if an image is funny or not? Can it point out the dog in the picture? The paper shows that AI is actually pretty good at this. It can often say, "Yes, that meme is funny," or "That is a politician." It's like a student who can spot the punchline in a joke book but doesn't know why it's funny.Level 2: Interpretation and Reasoning (The "Why")
This is the middle rung, and it's where things get messy. Can the AI explain why the joke works? Does it understand that the joke is making fun of a specific politician, or that it's using a cultural reference from 1990s TV? The paper finds that AI struggles here big time. It often gives answers that sound right but are actually wrong, or it misses the point entirely. It's like a student who can repeat a joke but explains it in a way that kills the humor because they missed the hidden meaning.Level 3: Generation (The "Make It")
This is the top rung. Can the AI create a funny meme or write a funny caption? The paper treats this as a new frontier. To make a good joke, you have to understand the rules of humor first. The researchers found that AI often tries to make jokes that are just... boring. They might be grammatically correct, but they aren't actually funny because the AI doesn't truly "get" the mechanism of humor.
The Three Big Problems
The authors identified three main reasons why AI is having such a hard time with jokes:
- The "Shortcut" Trap: AI models are smart, but they are also lazy. They often learn to guess the answer by looking for easy patterns. For example, if a picture has a politician's face and a sad cloud, the AI might guess "sad" or "funny" just because it saw that combination before, without actually understanding the story. The paper suggests that many current tests are too easy, letting the AI cheat by using these shortcuts instead of really thinking.
- The Cultural Blind Spot: Jokes rely on shared knowledge. If you don't know who a specific celebrity is, or what a recent news event was, you won't get the joke. The paper points out that AI is mostly trained on English and Western internet culture. It's like a student who only knows jokes from one country and is completely lost when someone tells a joke from a different culture. It misses the "inside jokes" that make humor work.
- The "Story" Problem: Many funny things, like comic strips, tell a story across several panels. The joke happens because of what happened in the first panel compared to the last one. The paper found that AI is often bad at connecting the dots between these panels. It looks at each picture separately instead of seeing the whole story, missing the timing and the surprise that makes a comic strip funny.
What the Paper Actually Found
The researchers didn't just guess; they tested many different AI models on a bunch of different joke datasets. They found that:
- AI is getting better at spotting funny things, but it's still far behind humans when it comes to explaining them.
- The best models (like some very large, expensive ones) can get about 80% of the recognition tasks right, but they drop significantly when asked to explain the joke or pick the right title for a cartoon.
- Humans are still the kings of comedy. In the tests where humans were compared to AI, humans consistently understood the jokes much better, especially when the jokes required understanding social norms or cultural references.
The paper also warns us about the future. As AI gets better at making jokes, it could accidentally make mean or harmful jokes because it doesn't understand the difference between "funny" and "hurtful." It's like giving a robot a stand-up comedy microphone without teaching it about empathy.
The Bottom Line
This paper is a reality check. It tells us that while our AI friends are becoming amazing at describing the world, they are still very confused about the meaning behind the world's jokes. They can see the dog, but they don't get why the dog is wearing a hat and acting like a human. To fix this, the authors say we need to stop just testing if AI can guess the right answer, and start testing if it can actually understand the story, the culture, and the reason behind the laugh. Until then, the AI might be able to tell a joke, but it probably won't know when to laugh at its own punchline.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.