← Latest papers
💻 computer science

MetaphorVU: Towards Metaphorical Video Understanding

This paper introduces MetaphorVU-Bench, the first comprehensive benchmark for metaphorical video understanding, reveals that current MLLMs struggle with cross-domain mapping, and proposes MetaphorBoost, an inference-time framework leveraging a metaphor knowledge graph to significantly enhance performance.

Original authors: Zhuoqun Li, Boxi Cao, Guiping Jiang, Fangrui Lv, Ruotong Pan, Jianan Wang, Xiangyu Wu, Hongyu Lin, Yaojie Lu, Yong Du, Ruyin Jia, Liyan, Tingting Gao, Han Li, Xianpei Han, Le Sun

Published 2026-05-26
📖 4 min read☕ Coffee break read

Original authors: Zhuoqun Li, Boxi Cao, Guiping Jiang, Fangrui Lv, Ruotong Pan, Jianan Wang, Xiangyu Wu, Hongyu Lin, Yaojie Lu, Yong Du, Ruyin Jia, Liyan, Tingting Gao, Han Li, Xianpei Han, Le Sun

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a short video. On the surface, you see a group of pigs wearing fancy tuxedos eating at a lavish banquet, while cats scavenge for scraps under the table.

A human viewer instantly understands the real story: This isn't about animals; it's a critique of a corrupt ruling class (the pigs) stealing wealth from the poor (the cats). This is metaphor: using one thing to represent something else entirely.

This paper, MetaphorVU, argues that while Artificial Intelligence (AI) has gotten very good at describing what is in a video (e.g., "There is a pig"), it is terrible at understanding what the video means (e.g., "This is a political satire").

Here is a simple breakdown of their findings and solution:

1. The Problem: AI is "Literal-Minded"

The researchers built a new test called MetaphorVU-Bench. Think of this as a "final exam" for AI, featuring 860 real-world videos that rely on metaphors (like the pig banquet).

They tested the smartest AI models available today (including giants like GPT-5 and Gemini).

  • The Result: The AI models failed miserably. They scored around 64 out of 100, while humans scored around 83.
  • The Diagnosis: The AI didn't fail because it couldn't see the pigs or the cats. It failed because it couldn't make the connection. It saw the "visual elements" but couldn't map them to the "underlying concepts."
    • Analogy: It's like a translator who knows every word in a dictionary but doesn't understand the jokes or idioms. They can translate "It's raining cats and dogs" literally, but they don't get that it just means "it's raining hard."

2. The Solution: Giving AI a "Cheat Sheet"

The researchers realized the AI wasn't stupid; it just lacked the specific "cultural knowledge" needed to make these leaps. To fix this, they built MetaphorBoost.

  • The Metaphor Knowledge Graph: Imagine a giant, digital web connecting concepts. In this web, "pig" is linked to "greed," "tuxedo" is linked to "power," and "banquet" is linked to "excess." This isn't just a list of facts; it's a map of how metaphors work.
  • How it Works: When the AI watches a video, instead of guessing the meaning on its own, MetaphorBoost pauses and asks this "Knowledge Graph": "Hey, I see a pig in a tuxedo. What does that usually symbolize in human culture?" The graph replies: "Power and greed."
  • The Result: With this "cheat sheet" (or external scaffold), the AI's performance jumped up significantly. It didn't just get better at guessing; it got better at reasoning because it had the right tools to connect the dots.

3. The 8 Types of Metaphors

The paper also organized these tricky videos into 8 categories, showing that metaphors come in many flavors:

  1. Body Language: Using exaggerated movements (like a student dancing then collapsing) to show hope turning to despair.
  2. Atmosphere: Using dark lighting or sad music to show loneliness.
  3. Cultural Symbols: Using a specific object (like a lantern) to represent a wish for success.
  4. Naturalistic Symbols: Using a wilting flower to represent a dying relationship.
  5. Causal Montage: Showing a cause and effect (e.g., putting on a ring \rightarrow sweeping floors) to imply "marriage brings hard work."
  6. Analogical Montage: Showing two similar things side-by-side (e.g., a childhood game and an adult job) to show how we miss our youth.
  7. Surreal Narrative: Using impossible things (like talking animals) to tell a story about real life.
  8. Performative Narrative: Using actors in a skit to criticize social behavior.

The Bottom Line

The paper concludes that current AI is great at perception (seeing the world) but weak at cognition (understanding the deeper meaning of the world).

To make AI truly understand human communication, we can't just make the models bigger; we have to give them better "maps" of how humans connect ideas. By adding a Metaphor Knowledge Graph, they showed that AI can learn to "get the joke" and understand the hidden meanings in our videos.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →