← Latest papers
💬 NLP

I Came, I Saw, I Explained: Benchmarking Multimodal LLMs on Figurative Meaning in Memes

This paper benchmarks eight state-of-the-art multimodal large language models on their ability to detect and explain figurative meanings in internet memes, revealing that while the models can identify such meanings, they exhibit a strong bias toward assuming figurative content exists and often fail to provide faithful explanations even when their predictions are correct.

Original authors: Shijia Zhou, Saif M. Mohammad, Barbara Plank, Diego Frassinelli

Published 2026-03-25
📖 5 min read🧠 Deep dive

Original authors: Shijia Zhou, Saif M. Mohammad, Barbara Plank, Diego Frassinelli

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet is a giant, chaotic party where everyone is shouting jokes, sharing memes, and making inside jokes. A meme is like a postcard from this party: it has a picture and some words, but the real joke often isn't in the words or the picture alone. It's in the twist between them.

For example, if a picture shows a cat looking grumpy and the text says, "I came, I saw, I complained," the joke is that it's mocking a famous heroic quote. You need to understand the history, the tone, and the visual cue to get it. This is called figurative meaning—saying one thing but meaning another, or using a picture to say something deeper.

This paper is like a report card for the newest, smartest AI robots (called Multimodal Large Language Models or MLLMs) to see if they can actually "get" these jokes.

Here is the breakdown of what the researchers did and what they found, using some everyday analogies:

1. The Test: "The Meme Olympics"

The researchers took eight of the smartest AI models available (think of them as the top athletes in the AI world) and put them through a grueling test.

  • The Athletes: Models from families like Aya, Gemma, and Qwen (ranging from small, nimble ones to massive, heavy-weight giants).
  • The Events: They showed the AI three different sets of memes.
  • The Tasks:
    1. Detect: "Is this meme making a joke (figurative), or is it just a normal picture (literal)?"
    2. Explain: "Why is it a joke? Tell me the story."

2. The Big Surprise: The "Over-Enthusiastic Detective"

The biggest finding was that the AI models are like conspiracy theorists or over-eager detectives.

  • The Bias: Even when a meme is totally serious and literal (like a photo of a dog eating a sandwich), the AI almost always says, "Ah! This is a deep metaphor! There is hidden meaning here!"
  • The Analogy: Imagine a detective who sees a man buying milk and immediately concludes, "He is definitely a spy hiding a secret code in the milk carton!" The AI just loves finding hidden meanings, even when there aren't any. It assumes every meme is a puzzle to be solved, even when it's just a picture.

3. The "Blindfold" Test: What Do They Need to See?

To figure out how the AI understands these jokes, the researchers played a trick. They showed the AI the memes in three ways:

  1. Full Meme: Picture + Text.
  2. Text Only: Just the words (the picture was erased).
  3. Image Only: Just the picture (the words were erased).

The Results:

  • Irony and Sarcasm: These are like visual punchlines. If you take away the picture, the AI gets confused. It needs to see the grumpy face or the weird situation to get the joke.
  • Metaphors: These are like complex recipes. You need both the ingredients (text) and the cooking method (image) to make the dish. If you give the AI just the text or just the image, it fails miserably. It needs the whole package to understand the connection.
  • Exaggeration: Interestingly, sometimes the AI got better at spotting exaggeration when the text was removed, suggesting it was relying too much on the words before and getting distracted.

4. The "Hallucination" Problem: Lying with Confidence

The researchers asked the AI to explain why it thought a meme was funny. This is where things got messy.

  • The Problem: Even when the AI guessed the right label (e.g., "This is sarcasm"), its explanation was often made up.
  • The Analogy: Imagine a student taking a test. They guess the right answer, but when the teacher asks, "Show your work," the student writes, "Because the cat in the picture is wearing a hat, and hats mean sarcasm."
    • The answer was right, but the reasoning was nonsense.
    • The AI often "hallucinates" details that aren't there (like seeing a finger over a cat's mouth when it's actually lifting the nose) or misses the real human emotion behind the joke.

5. The "Human Touch" Gap

The researchers also had humans grade the AI's explanations.

  • The Verdict: The AI is great at reading the text but terrible at reading the room. It misses the social nuance.
  • Example: One meme showed someone willing to spend $1,000 on a phone but getting mad about a 99-cent app. The joke is about human psychology (we rationalize big spending but hate small fees). Half the AI models completely missed this, treating it as a literal math problem instead of a social observation. They lack "street smarts."

The Bottom Line

This paper tells us that while AI is getting incredibly good at reading and seeing, it still struggles with the soul of a meme.

  • It's too eager: It sees metaphors where there are none.
  • It's a bad storyteller: It often makes up reasons for its guesses.
  • It lacks life experience: It doesn't understand the weird, irrational ways humans behave, which is often what makes a meme funny.

In short: The AI is like a brilliant robot who has read every book in the library but has never actually gone to a party. It knows the rules of language, but it hasn't quite learned the art of the inside joke yet.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →