Beyond Understanding: Evaluating the Pragmatic Gap in LLMs' Cultural Processing of Figurative Language
This paper evaluates the pragmatic gap in large language models' cultural processing by demonstrating that while they can interpret figurative language, they struggle significantly more with culturally grounded expressions like Egyptian Arabic idioms and proverbs compared to English, prompting the release of the Kinayat dataset to support future research.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a brilliant new student named LLM (Large Language Model). This student has read almost every book, website, and article on the internet. They can recite facts, write poems, and solve math problems better than almost anyone.
But there's a catch: LLM is like a tourist who has memorized a phrasebook but hasn't actually lived in the country. They know the words, but they don't quite get the vibe, the inside jokes, or the unwritten rules of how people actually talk to each other.
This paper is like a "cultural pop quiz" designed to see if our super-smart student can actually get the local culture, or if they are just faking it.
The Test: Idioms and Proverbs
The researchers didn't ask LLM to solve calculus. Instead, they tested it on idioms and proverbs.
Think of an idiom like a secret handshake.
- Literal meaning: "Don't sell water in the village of water-sellers."
- Real meaning: "Don't try to sell your wares to people who already have plenty of them" (or, "Don't try to outsmart the experts").
If you just translate the words, it makes no sense. You need to understand the story behind it. The paper tested LLMs on:
- English Proverbs: The "textbook" examples.
- Arabic Proverbs: The "standard" cultural examples.
- Egyptian Arabic Idioms: The "street smart," slang-heavy, local jokes.
The Results: The "Tourist" vs. The "Local"
The results revealed a clear hierarchy, like a video game with increasing difficulty levels:
Level 1: English Proverbs (The Easy Mode)
LLMs were pretty good here. They could explain the meaning of English sayings because they had read so many English books. They got about 90% right.Level 2: Arabic Proverbs (The Medium Mode)
When they switched to formal Arabic proverbs, the scores dropped by about 4%. The models were still okay, but they started to stumble a bit on the cultural nuances.Level 3: Egyptian Idioms (The Hard Mode)
This is where it got messy. When the models tried to understand Egyptian slang and idioms (which are very specific to daily life, work, and family dynamics), their scores dropped another 10%. They were essentially guessing.
The Big Surprise: The "Pragmatic Gap"
Here is the most interesting part of the story. The researchers gave the models two types of tests:
- Test A (The Quiz): "Here is an idiom. What does it mean?"
- Result: The models were decent at this. They could pick the right definition from a list.
- Test B (The Party): "Here is a situation. Use an idiom to describe it naturally."
- Result: Total failure.
The Analogy:
Imagine you are at a party. You know the definition of the word "awkward."
- Test A: If someone asks, "What does 'awkward' mean?" you can say, "It means uncomfortable." (You pass).
- Test B: If someone drops a tray of glasses, and you immediately shout, "That is a classic example of awkward!" while everyone is staring at you, you have failed the party. You know the word, but you don't know when or how to use it.
The paper found that LLMs are terrible at Test B. Even the smartest models dropped their accuracy by 14% when asked to use the idiom correctly in a sentence, compared to just explaining it. They sound like a robot reading a dictionary, not a human having a conversation.
The "Cultural Blind Spot"
The study also looked at connotations (the emotional feeling of a word).
- Is a phrase funny? Is it insulting? Is it sarcastic?
- Even when humans agreed 100% on the meaning, the models only agreed with them about 85% of the time.
It's like the models are colorblind to the "emotional colors" of language. They see the black and white text, but they miss the rainbow of feelings behind it.
The New Tool: "Kinayat"
To help fix this, the authors created a new dataset called Kinayat.
Think of this as a new "Cultural Survival Guide" for AI. It's a collection of 325 Egyptian idioms, carefully labeled with their meanings and examples of how to use them in real life. It's like giving the tourist a map that doesn't just show the streets, but also shows where the good coffee shops are and which streets to avoid.
The Bottom Line
This paper teaches us that knowing a language is not the same as understanding a culture.
Current AI models are like encyclopedias that can recite facts perfectly. But when it comes to the messy, emotional, cultural, and "inside joke" parts of human conversation, they are still just tourists. They can read the menu, but they haven't learned how to order the food like a local yet.
The takeaway: If you want an AI to truly understand your culture, you can't just feed it more text. You have to teach it the context, the feelings, and the unwritten rules of how people actually talk.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.