← Latest papers
💬 NLP

On the Wings of Imagination: Conflicting Script-based Multi-role Framework for Humor Caption Generation

This paper introduces HOMER, a novel humor-theory-driven multi-role LLM framework that leverages script oppositions, retrieval-augmented hierarchical imagination, and targeted caption generation to significantly outperform existing state-of-the-art methods in creating funny multi-modal image captions.

Original authors: Wenbo Shang, Yuxi Sun, Jing Ma, Xin Huang

Published 2026-08-04
📖 4 min read☕ Coffee break read

Original authors: Wenbo Shang, Yuxi Sun, Jing Ma, Xin Huang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to tell a joke. You might think, "Just give it a picture and ask it to be funny!" But here's the catch: for a computer, "funny" is a slippery, confusing concept. It's not just about describing what you see; it's about understanding the hidden rules of the world, spotting when those rules are broken, and then twisting your brain to find the surprise in that break. This is the heart of multimodal humor generation, a field where scientists try to help Artificial Intelligence (AI) understand the difference between a normal coffee cup and a coffee cup the size of a bathtub.

To do this, the researchers in this paper rely on a classic idea from humor theory called Script Opposition. Think of a "script" as a mental movie script we all know by heart, like "going to a meeting" or "drinking coffee." Humor often happens when two of these scripts crash into each other in a weird way—like a serious business meeting where everyone is drinking from giant mugs. The paper suggests that to make a computer truly funny, it can't just guess; it needs a structured plan to find these crashes, imagine wild connections, and then write a punchline that ties it all together.

The paper introduces a new system called HOMER (Humor-theory-driven Multi-role LLM collaboration framework with humor retrieval). Instead of asking a single AI to "be funny" and hoping for the best, HOMER acts like a tiny, specialized comedy writing team. It splits the job into three distinct roles, each played by a different AI agent working together:

  1. The Conflict Detective (Conflicting-script Extractor): This agent looks at the image and the scene description to find the "glitch" in reality. If the image shows a meeting room with normal chairs but one giant, oversized coffee cup, this agent spots the clash between "normal office" and "giant cup." It identifies the core contradiction that makes the joke possible.
  2. The Wild Imaginator (Hierarchical Imaginator): Once the conflict is found, this agent goes on a creative hunt. It takes the weird object (the giant cup) and builds a tree of associations. It might think: Cup → Coffee → Milk → Cow. But it doesn't just guess; it checks a massive database of real human jokes to see which connections actually land as funny. It prunes away the boring ideas and keeps the ones that have the most "humor potential."
  3. The Punchline Writer (Caption Generator): Finally, this agent takes the conflict, the wild imagination tree, and the scene description to write the actual caption. It combines all the clues into a short, witty sentence that explains the absurdity.

The paper finds that this team approach works significantly better than letting a single AI try to do everything at once. When tested on thousands of real-world cartoon captions from the New Yorker (a famous magazine known for its funny drawings), HOMER consistently outperformed other top AI models. In fact, on average, HOMER's captions were about 7% more likely to win against human-written captions compared to the next best method.

The authors argue against the idea that simply making an AI "think harder" or "reason more" (like using complex step-by-step logic chains) is enough to generate humor. They show that without a specific guide on how humor works (like looking for script oppositions) and without a structured way to explore creative ideas, AI tends to produce captions that are logical but not actually funny. By breaking the process down into these specific, theory-driven steps, HOMER creates captions that feel more original and genuinely amusing, proving that sometimes, to be funny, you need a little bit of structure and a lot of imagination.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →