Unmasking the Factual-Conceptual Gap in Persian Language Models
This paper introduces DivanBench, a diagnostic benchmark for Persian language models that reveals a significant gap between retrieving cultural facts and reasoning about implicit social norms, demonstrating that current models suffer from acquiescence bias and fail to internalize cultural schemas despite continuous pretraining.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to understand human culture. You give it a library containing billions of books, articles, and stories written in Persian. You ask the robot: "Do you understand how Iranians behave?"
The robot says, "Yes! I know all the words. I know that Iranians say 'Taarof' (polite refusal) when offered food. I know they burn Esfand seeds to ward off bad luck."
But then, you put the robot in a real-life situation. You offer a guest food, and they refuse. You offer again, they refuse. You offer a third time, and they finally accept. The robot, however, gets confused. It thinks, "Wait, the first time they refused, that was a violation! They should have accepted immediately!" Or, it sees a guest who doesn't refuse and thinks, "That's perfectly normal!"
This is exactly what the paper "Unmasking the Factual-Conceptual Gap in Persian Language Models" is about. The researchers built a test called DIVANBENCH (named after a traditional Persian low table) to see if AI models actually understand Persian culture or if they are just parroting it.
Here is the breakdown of their findings using simple analogies:
1. The "Yes-Man" Problem (Acquiescence Bias)
The Analogy: Imagine a student taking a test. When the teacher asks, "Is it polite to stand up when an elder enters?" the student says "Yes!" correctly. But when the teacher asks, "Is it polite to sit and ignore the elder?" the student also says "Yes!" because they are too eager to agree with anything that sounds like a cultural rule.
The Finding: Most of the AI models tested were terrible "Yes-Men." They could identify when something was correct (84–92% accuracy), but they were terrible at spotting when something was wrong (only 19–48% accuracy). They were just pattern-matching keywords like "elder" or "food" and saying "Yes, that sounds Persian," without actually thinking about the logic.
2. The "More Data, Less Smarts" Paradox
The Analogy: Think of the base AI model (Llama 3.1) as a smart foreigner who has studied Persian grammar and culture from a textbook. They are skeptical and careful. Now, imagine you take that same person and force them to live in Iran for a year, reading only Persian social media and news (Continuous Pretraining).
You might expect them to become a cultural genius. Instead, they became a mindless conformist. They learned to repeat cultural phrases perfectly, but they lost their ability to question them.
The Finding: When the researchers took a standard model and gave it extra Persian training (creating a model called Dorna2), the model got worse at logic. Its ability to spot cultural violations dropped by 43%. It became fluent in the "sound" of the culture but lost the "sense" of it. It learned to mimic the vibe without understanding the rules.
3. The "Fact vs. Feeling" Gap
The Analogy: Imagine you have a recipe book for a complex dish.
- Factual Knowledge: You can recite the ingredients perfectly. "I know that Sard (cold) foods are bad for a fever."
- Conceptual Reasoning: You are in a kitchen with a sick friend. You see a melon. Do you give it to them?
- A robot with only "Factual Knowledge" might say, "Melons have vitamins! Give it to them!"
- A culturally competent human knows: "No! In Persian medicine, melons are Sard (cold). Giving a cold food to a sick person is a big no-no."
The Finding: The AI models were great at the recipe book (Factual MCQs). But when asked to apply that knowledge to a real-life scenario (Scenario-Based MCQs), their performance dropped by an average of 21%. They knew the facts, but they couldn't use them to navigate the social world.
4. Bigger Isn't Always Better
The Analogy: You might think a bigger library (a larger AI model) means a smarter librarian.
The Finding: The biggest model tested (Gemma3-12B) was the best at memorizing facts, but it was also one of the worst at spotting cultural violations. It was like a giant encyclopedia that knew every rule but couldn't tell you which rule applied to this specific moment. Bigger models just got better at memorizing the "pattern" of culture, not the "logic" of it.
The Big Takeaway
The paper concludes that culture is not just a list of facts to be memorized. It is a complex, invisible set of rules about when to do things, how to do them, and why they matter.
Current AI models are like tourists who have memorized a phrasebook. They can say the right words, but if you put them in a tricky social situation, they will likely offend everyone because they don't understand the underlying "schema" (the mental map) of how the culture actually works.
In short: To build AI that truly understands a culture, we can't just feed it more text. We need to teach it how to think like a human, not just how to sound like one.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.