FFE-Hallu:Hallucinations in Fixed Figurative Expressions:Benchmark of Idioms and Proverbs in the Persian Language
This paper introduces FFE-Hallu, the first benchmark for evaluating hallucinations in fixed figurative expressions within the Persian language, revealing that current large language models struggle with cultural grounding and frequently generate or fail to detect fabricated idioms and proverbs.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to speak a human language. You might think that if the robot knows the dictionary definitions of words like "kick" and "bucket," it can understand the phrase "kick the bucket." But in reality, that phrase doesn't mean literally kicking a pail; it's a fixed, cultural code for "dying."
This paper, FFE-HALLU, is like a specialized "lie detector test" for AI, designed to see if these robots are actually making up their own cultural codes or if they truly understand the real ones.
Here is the breakdown of what the researchers did, using simple analogies:
1. The Problem: The "Fake Idiom" Trap
Large Language Models (LLMs) are great at writing sentences that sound right. But when it comes to Fixed Figurative Expressions (FFEs)—like idioms ("break a leg") and proverbs ("a stitch in time")—they often get confused.
The researchers define a specific type of error called Figurative Hallucination.
- The Analogy: Imagine a robot trying to tell a joke. It says, "Why did the scarecrow win an award? Because he was outstanding in his field!" That's a real joke.
- The Hallucination: The robot then says, "Why did the toaster win an award? Because he was crispy in his field!"
- It sounds like a joke. It follows the rules of grammar. But it's fake. It doesn't exist in human culture.
- The danger is that if a non-native speaker hears this, they might think, "Oh, that's a real Persian saying!" and use it, looking foolish or spreading misinformation.
2. The Solution: The "FFE-HALLU" Exam
The authors created a test called FFE-HALLU specifically for the Persian language. They built a "gym" with 600 exercises to test the AI's cultural muscle. The test has three main stations:
Station 1: The "Meaning-to-Phrase" Challenge
- The Task: The AI is given a definition (e.g., "To regret a mistake and promise not to do it again") and asked to say the actual Persian idiom for it.
- The Trap: Instead of saying the real idiom, the AI might invent a fake one that sounds plausible but has never been used by a human.
- The Result: Some top models (like GPT-4.1) got about 60% right. Others (like Qwen) were terrible, inventing fake idioms in nearly 70% of cases.
Station 2: The "Fake or Real?" Detective Game
- The Task: The AI is shown a phrase and asked, "Is this a real Persian idiom?"
- The Trap: The researchers created 200 "fake" idioms that look and sound exactly like real ones (e.g., swapping one word in a real phrase or twisting the meaning).
- The Result: This was the hardest part. Even smart models got confused. When asked "Is this fake?", many models said "No" (thinking it was real) when it was actually a lie. It's like a security guard who thinks a perfect fake ID is real because it looks so good.
Station 3: The "Translator" Test
- The Task: The AI is given an English idiom (e.g., "It's raining cats and dogs") and must find the equivalent Persian idiom, not a word-for-word translation.
- The Trap: If the AI translates it literally ("It is raining felines and canines"), that is a hallucination because it's not how Persian speakers talk.
- The Result: Most models struggled to find the cultural equivalent, often inventing new phrases or translating too literally.
3. The Findings: Who Passed and Who Failed?
The researchers tested six different AI models:
- The Top Performers: GPT-4.1 and Claude 3.7 were the best. They could usually find the real idioms and reject the fake ones, though they still made some mistakes.
- The Strugglers: Models like DeepSeek, Gemma, and Qwen often failed. They tended to "over-generate," meaning they confidently made up fake idioms that sounded very real but were completely invented.
- The "Yes-Man" Problem: Some models were so eager to please that when asked "Is this fake?", they would say "No" (meaning "Yes, it's real") even when it was a complete fabrication.
4. The "Judge" Problem
The paper also asked: Can an AI judge another AI's work?
They tried using one AI to grade the answers of the others.
- The Analogy: It's like asking a student to grade another student's essay.
- The Result: It worked okay for easy questions, but when it came to spotting subtle "fake idioms," the AI judges often missed them. They couldn't tell the difference between a "wrong answer" and a "made-up answer." The researchers found that human experts (native Persian speakers) were still the only ones who could reliably catch these subtle lies.
Summary
This paper is a warning label for AI. It shows that while AI can write fluent sentences, it often lacks cultural grounding. It can sound like it knows a culture's idioms, but it's actually just making them up on the spot.
The researchers built a test to prove this, showing that even the smartest AI models are prone to "figurative hallucinations"—telling cultural lies that sound like the truth. Until AI learns to distinguish between a real cultural saying and a convincing fake, we can't fully trust them with tasks that require deep cultural understanding.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.