← Latest papers
💻 computer science

MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes

This paper introduces MemeBench, a diagnostic benchmark utilizing the VIKR schema to evaluate and reveal the specific cultural knowledge gaps in Large Vision-Language Models' interpretation of memes, demonstrating that while entity-guided retrieval improves knowledge integration, it currently trades off visual coverage.

Original authors: Weihang Wang, Kainan Tu, Jielei Zhang, Run Yang, Boheng Sheng, Yuchen He, Yu Xie, Pengyu Chen, Peiyi Li, Huyang Sun, Longwen Gao, Zhouhui Lian

Published 2026-07-31
📖 4 min read☕ Coffee break read

Original authors: Weihang Wang, Kainan Tu, Jielei Zhang, Run Yang, Boheng Sheng, Yuchen He, Yu Xie, Pengyu Chen, Peiyi Li, Huyang Sun, Longwen Gao, Zhouhui Lian

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking into a room full of people speaking a language you don't know, but you can see their faces, their clothes, and the objects they are holding. You are a super-smart robot designed to describe exactly what you see. You can say, "That person is wearing a red hat," or "They are holding a golden cup." You are perfect at describing the pixels. But then, someone points to a picture of a cat wearing a tiny tuxedo and says, "This is a meme about how cats think they own the internet." If you don't know the internet, if you don't know the culture of cat owners, or if you don't understand the joke, you might just say, "It is a cat in a suit." You described the picture perfectly, but you completely missed the point.

This is the gap that researchers are trying to bridge with Large Vision-Language Models (LVLMs). These are AI systems that can both "see" images and "speak" about them. For a long time, we thought that if an AI could describe an image accurately, it understood it. But this paper argues that description is not the same as interpretation. It's the difference between reading the words of a joke and actually getting why it's funny. To understand a meme, you need more than just your eyes; you need a library of cultural knowledge, inside jokes, and community history that isn't written on the image itself. If an AI can't access that hidden library, it's like a tourist who can read a map but doesn't know the local slang or the history of the landmarks.

Enter MemeBench, a new diagnostic tool created by researchers from Bilibili, Fudan University, and Peking University to test exactly where these AI models get stuck. Think of MemeBench as a "cultural pop quiz" for AI, specifically focused on the wild, fast-moving world of anime, comics, games, and internet subcultures (often called ACG). The researchers collected 1,253 memes in both Chinese and English. But instead of just asking the AI, "Is this funny?" or "What is the answer?", they asked the AI to write a full explanation. Then, they broke those explanations down into four specific ingredients using a system they call VIKR:

  1. Visual: Did you see what's actually in the picture?
  2. Identity: Did you know who or what the characters are? (e.g., Is that a generic ninja, or is it specifically Naruto?)
  3. Knowledge: Did you know the backstory or the cultural rule that makes this funny?
  4. Reasoning: Did you connect the dots to explain the joke?

The researchers tested 26 different AI models, from the biggest commercial ones to open-source experiments. Here is the big surprise they found: Every single model was better at describing the picture than at understanding the joke. Even the smartest AI in the test could describe the visual details with high accuracy, but when it came to the cultural knowledge needed to interpret the meme, it stumbled. The best model still had a "gap" of 22.6% between how well it saw the image and how well it knew the culture. It's like having a pair of eyes that work perfectly, but a brain that's missing the instruction manual for the world.

The paper also argues against the idea that simply giving the AI more "reasoning steps" (telling it to "think step-by-step") fixes this problem. It doesn't. The models still missed the cultural context. However, the researchers did find a way to help. They introduced a method called KAR (Knowledge-Aware Retrieval). Imagine if, before the AI tried to tell the joke, it was allowed to quickly look up the specific character names and history in a specialized encyclopedia (which they built called CultureBase) before searching the general web. This didn't just make the AI smarter; it made it more precise. It helped the AI find the missing cultural pieces (Identity and Knowledge) without messing up the visual description.

In short, this paper suggests that while AI is getting great at being a camera, it is still struggling to be a cultural critic. The models aren't failing because they can't see; they are failing because they don't know the story behind the image. By using MemeBench, researchers can now pinpoint exactly which part of the explanation is missing—whether it's the name of the character, the history of the event, or the logic of the joke—rather than just giving a single score. This helps developers build AI that doesn't just describe the world, but actually understands the culture within it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →