← Latest papers
💬 NLP

MasalBench: A Benchmark for Contextual and Cross-Cultural Understanding of Persian Proverbs in LLMs

This paper introduces MasalBench, a benchmark for evaluating the contextual and cross-cultural understanding of Persian proverbs in large language models, revealing that while these models excel at identifying proverbs within their native context, they struggle significantly with finding equivalent English proverbs, thus highlighting limitations in their cultural knowledge and analogical reasoning capabilities.

Original authors: Ghazal Kalhor, Behnam Bahrak

Published 2026-01-30
📖 4 min read☕ Coffee break read

Original authors: Ghazal Kalhor, Behnam Bahrak

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to understand human conversation. You want to see if this robot can truly "get" the jokes, idioms, and old sayings people use every day, or if it just sounds like a dictionary that has memorized definitions.

This paper introduces a new test called MasalBench to see how well AI models (the "robots") understand Persian proverbs. Think of Persian proverbs as the "spice" of the language—short, colorful sayings that carry deep wisdom, like "Don't count your chickens before they hatch," but specific to Iranian culture.

Here is the breakdown of what the researchers did and what they found, using some simple analogies:

1. The Problem: The "Low-Resource" Gap

Most AI models are like students who have read millions of English books. They are experts in English. But for languages like Persian, the AI has read far fewer books. The researchers wanted to know: Can these smart robots still understand the "spice" of a language they haven't studied as much?

2. The Test: MasalBench (The "Proverb Gym")

The team built a gym with two different types of exercises to test the AI's muscles:

  • Exercise A: Contextual Understanding (The "Conversation Test")

    • The Setup: They created 1,000 short stories (dialogues) where two people are talking. One person uses a Persian proverb, and the AI has to guess why they said it.
    • The Trap: To make it hard, they didn't just give the right answer. They gave three wrong answers:
      1. The Literal Trap: Taking the words too seriously (e.g., thinking a proverb about "burning your mouth" is actually about hot soup).
      2. The Plausible Trap: A guess that sounds smart and logical but is actually wrong.
      3. The Irrelevant Trap: A guess that has nothing to do with the story.
    • The Result: The AI models were very good at this. They scored over 90%. It's like a student who has studied hard and can easily understand a joke when they hear it in a conversation.
  • Exercise B: Cross-Cultural Understanding (The "Translation Puzzle")

    • The Setup: They gave the AI a Persian proverb and asked: "Which English proverb means the same thing?"
    • The Challenge: This is like asking someone to find a twin in a different country. The "twins" might look different (different words or images) but have the same soul.
    • The Result: The AI struggled here. Their scores dropped significantly (the best one got about 79%). It's like a student who can understand a joke in English but gets confused when asked to find the equivalent joke in French. They know the words, but they miss the cultural "vibe."

3. The Big Takeaways

  • Size Matters, But Not Everything: Bigger AI models generally did better, but not always. Sometimes, a model that was specifically "trained to follow instructions" did better than a massive model that was just "trained to think."
  • The "Smart" Mistake: When the AI got the conversation test wrong, it usually picked the "Plausible Trap." This means the AI was trying to be logical and make sense of the story, but it missed the specific cultural nuance. It was thinking too hard and over-analyzing, rather than just "getting" the feeling.
  • The Cultural Wall: The biggest hurdle was connecting Persian wisdom to English wisdom. The AI is great at understanding the language it's in, but it's not great at jumping across cultural bridges to find the "spiritual twin" of a saying in another language.

In a Nutshell

The paper concludes that while modern AI is getting very good at understanding how Persian proverbs are used in daily chat, it still has a hard time understanding what those proverbs mean when you try to translate their spirit into English. It's a bit like a person who can perfectly mimic an accent but still doesn't fully understand the local humor.

The researchers made this test available for others to use, hoping it will help build better AI that understands not just words, but the culture behind them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →