Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages
This paper introduces MIDI, a new multilingual dataset spanning high-, medium-, and low-resource languages that embeds idioms in sentence and conversational contexts to reveal that current models struggle significantly with literal interpretations and low-resource languages, even when provided with contextual cues.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to understand human conversation. You give it a dictionary and thousands of books, and it learns to speak many languages fluently. But then, you ask it a simple question: "What does it mean when someone says they are 'raining cats and dogs'?"
If the robot answers, "It is literally raining animals," it has failed. It missed the idiom—a phrase where the words mean something totally different than their literal definition.
This paper introduces a new tool called MIDI (Multilingual Idiom Dataset) to test how well modern AI models handle these tricky phrases across 18 different languages, ranging from major global languages (like Chinese and Japanese) to smaller, less-digital languages (like Yoruba and Javanese).
Here is a breakdown of what the researchers found, using simple analogies:
1. The Problem: The "Literal Trap"
Idioms are like cultural inside jokes. To understand them, you can't just look up the words; you need to know the "vibe" of the culture.
- The Study: The researchers built a massive test bank (MIDI) with over 2,000 idioms. They didn't just give the AI the phrase; they put it in two different settings:
- The Sentence: A single line of text (like a flashcard).
- The Conversation: A short chat between two people (like a scene from a movie).
- The Goal: To see if the AI can tell the difference between a phrase used figuratively (the joke meaning) and literally (the actual meaning).
2. The Results: The "Rich vs. Poor" Gap
The researchers tested the smartest AI models available (both the expensive, "proprietary" ones and the free, "open-source" ones).
- The Language Divide: Think of language resources like internet bandwidth.
- High-Resource Languages (e.g., English, Chinese): These languages have huge amounts of data on the internet. The AI models performed decently here, like a student who has read a lot of textbooks.
- Low-Resource Languages (e.g., Yoruba, Minangkabau): These languages have very little data online. The models struggled badly here, often performing no better than random guessing. It's like asking a student to take a test in a language they've never seen before.
- The Literal Struggle: Even the best models found literal meanings much harder than figurative ones.
- Analogy: If you say, "I'm breaking a leg," the AI is great at knowing you mean "good luck" (figurative). But if you say, "I am breaking a leg" while actually in a hospital, the AI often gets confused and still thinks you mean "good luck." It has a hard time realizing when the joke is not a joke.
3. The "Memory" vs. "Reasoning" Test
The researchers wanted to know: Is the AI actually thinking, or is it just memorizing answers like a parrot?
- The Memory Test: They asked the AI, "What does this phrase mean?" with no context.
- Result: The big, expensive models were great at this. They had "memorized" the definitions.
- The Reasoning Test: They gave the AI the phrase plus the English definition and asked it to pick the right meaning based on the story.
- Result: This helped the open-source models a lot, but the gap between "rich" and "poor" languages remained.
- The "Bias" Discovery: The AI has a strong bias toward jokes. When in doubt, it assumes the phrase is figurative.
- Analogy: It's like a person who always assumes you are making a joke, even when you are being serious. This makes them good at spotting idioms but terrible at spotting when someone is being literal.
4. The "Steering" Experiment (The Remote Control)
The researchers tried a cool trick called Activation Steering. Imagine the AI's brain is a radio, and they found a specific "knob" that controls whether the AI is thinking about "memorized facts" or "logical reasoning."
- What they did: They turned this knob slightly to push the AI toward better reasoning.
- The Result: It worked! The AI got slightly better, especially for the low-resource languages. It's like giving a student a tiny hint or a nudge in the right direction, which helped them solve the puzzle they were stuck on. Interestingly, they found that a "knob" trained on a completely different dataset (MMLU-Pro) worked just as well on this idiom test, suggesting that the way AI handles memory and reasoning is a universal "muscle" across different tasks.
5. Humans vs. Machines
Finally, they asked real humans to take the same test.
- The Score: Humans got 94% correct.
- The AI Score: The best AI models got around 81% (proprietary) and 72% (open-source).
- The Gap: The biggest gap was in low-resource languages. For example, in Yoruba, the best open-source model scored only 23%, while humans scored 96%.
Summary
The paper concludes that while AI is getting very good at languages, it still struggles with the "soul" of language: idioms.
- It relies too much on memorization.
- It gets confused when a phrase is meant to be taken literally.
- It performs poorly in languages where it hasn't seen enough data.
- Even the smartest models are still far behind humans in understanding the nuance of cultural expressions.
The researchers didn't claim this will immediately fix translation apps or create new medical tools. Instead, they simply built a better "gym" (the MIDI dataset) to show us exactly where AI is weak, so we can build stronger models in the future.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.