← Latest papers
💬 NLP

When Meaning Isn't Literal: Exploring Idiomatic Meaning Across Languages and Modalities

This paper introduces "Mediom," a multilingual multimodal corpus of 3,533 idioms, and "HIDE," a hint-based framework, to address and improve the systematic failures of current AI models in understanding culturally grounded, non-literal idiomatic reasoning across text and image modalities.

Original authors: Sarmistha Das, Shreyas Guha, Suvrayan Bandyopadhyay, Salisa Phosit, Kitsuchart Pasupa, Sriparna Saha

Published 2026-04-14
📖 4 min read☕ Coffee break read

Original authors: Sarmistha Das, Shreyas Guha, Suvrayan Bandyopadhyay, Salisa Phosit, Kitsuchart Pasupa, Sriparna Saha

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to explain a joke to someone who speaks a different language. You say, "It's raining cats and dogs!" They look at the sky, confused, and ask, "Where are the animals? Is this a zoo?"

This is exactly the problem modern AI faces with idioms.

Idioms are phrases where the literal meaning is completely different from the real meaning. "Angur fol tok" in Bengali literally means "grapes are sour," but it actually means "sour grapes"—pretending you don't want something because you can't have it. If an AI just looks at the words, it sees fruit. If it understands the culture, it sees human psychology.

This paper introduces a new toolkit to teach AI how to stop taking things literally and start understanding the "human" side of language. Here is the breakdown in simple terms:

1. The Problem: AI is a Literal Robot

Current AI models (like the ones powering chatbots) are incredibly smart at grammar and facts. But when it comes to idioms, metaphors, and cultural sayings, they often fail. They are like a tourist who memorized a phrasebook but doesn't understand the local customs.

  • The Issue: If you show an AI a picture of a fox looking at grapes, it might describe the animal and the fruit. It misses the moral of the story (that the fox is just making excuses).
  • The Gap: Most research focuses on English. This paper looked at Hindi, Bengali, and Thai, languages rich in cultural metaphors that AI has largely ignored.

2. The Solution Part 1: "Mediom" (The New Textbook)

The researchers built a massive new library called Mediom.

  • What is it? Think of it as a bilingual picture dictionary for idioms. It contains 3,533 idioms from Hindi, Bengali, and Thai.
  • The Twist: For every idiom, they didn't just write a definition. They created:
    • A Gold-Standard Explanation: The perfect, human-written meaning.
    • A Matching Image: A picture that captures the feeling of the idiom, not just the literal words.
    • Cross-Language Links: Showing how similar ideas exist in different cultures (e.g., how "walls have ears" exists in all three languages).

Analogy: Imagine trying to learn to drive. Before, you only had a manual with text. Now, Mediom is a driving simulator that shows you the road, the traffic signs, and the feel of the car, all in three different countries.

3. The Solution Part 2: "HIDE" (The AI Coach)

Even with a great textbook, AI sometimes still gets the answer wrong. So, the researchers created a framework called HIDE (Hinting-based Idiom Explanation).

How HIDE works:

  1. The Mistake: The AI tries to explain an idiom and gets it wrong (e.g., it thinks "hitting with a shoe while smiling" is just about shoes).
  2. The Feedback Loop: The system catches this error. It doesn't just say "Wrong." It asks, "Why did you get this wrong? Was it too literal?"
  3. The Hint: The system generates a "Hint" (a tiny piece of advice) like, "Remember, this isn't about the shoe; it's about fake politeness."
  4. The Retry: The AI tries again, using that hint. It learns from its own mistake, just like a student correcting a math problem after seeing the teacher's note.

Analogy: Imagine playing a video game where you keep hitting a wall. Instead of just restarting, a coach whispers, "Don't run straight; jump over the gap." You try again, and this time, you win. HIDE is that whispering coach for AI.

4. The Results: Did it Work?

The researchers tested this on various AI models:

  • Text-Only AI (LLMs): These are like people who read a lot but have never seen a picture. They were okay at idioms, but HIDE made them much better, helping them understand the cultural "vibe."
  • Image + Text AI (VLMs): These are like people who can see the picture but struggle with the deep meaning. They were surprisingly bad at idioms at first, often getting confused by the literal images. However, with the Mediom dataset and HIDE, they started to understand the connection between the picture and the hidden meaning.

The Big Takeaway

This paper is a wake-up call. It says: "To make AI truly human-like, we can't just teach it words; we have to teach it culture, metaphors, and the ability to learn from its own mistakes."

By building a specialized library (Mediom) and a smart correction system (HIDE), the researchers are helping AI move from being a "dictionary that reads words" to a "friend that understands jokes."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →