TARAZ: Persian Short-Answer Question Benchmark for Cultural Evaluation of Language Models
This paper introduces TARAZ, a novel Persian short-answer benchmark and hybrid evaluation framework that overcomes the limitations of existing multiple-choice and exact-match metrics to more accurately assess the cultural competence of large language models through morphological normalization and semantic similarity scoring.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to understand human culture. You give it a test, but the test is written in a language the robot knows well (English), and the questions are simple "Yes/No" or "Multiple Choice" formats. The robot might get a high score just by memorizing facts, but it doesn't actually get the vibe, the humor, or the unwritten rules of the culture.
This paper, TARAZ, is like a new, much harder, and much fairer test designed specifically for Persian (Farsi) speakers to see if AI models truly understand their culture.
Here is the breakdown of what the researchers did, using some everyday analogies:
1. The Problem: The "Multiple Choice" Trap
Imagine you are testing a student's knowledge of American football.
- Old Method (Multiple Choice): You ask, "Who won the Super Bowl in 2020?" The student guesses "A." They get it right. But do they understand the game? No. They just memorized a fact.
- The Issue with Persian: Most AI tests for Persian culture were like this. They used Multiple Choice Questions (MCQs) and checked if the answer was an exact string match.
- The Persian Twist: Persian is a "shape-shifting" language. The word for "bread" can be written formally, informally, with different accents, or using different numbers (Arabic vs. Persian digits). If a robot says "bread" and the answer key says "bread (informal)," a strict computer says, "Wrong!" even though the meaning is identical. The old tests were like a teacher failing a student for spelling "color" as "colour."
2. The Solution: The "Short Answer" Conversation
The researchers created TARAZ, a new framework that treats the AI like a human in a conversation, not a machine filling out a bubble sheet.
- The Test Format: Instead of picking A, B, C, or D, the AI has to write a free-form answer.
- Example: "What do people usually do at an Iranian airport before a trip?"
- AI Answer: "They ask for forgiveness."
- Old Test: Would fail if it didn't match the exact phrase in the database.
- TARAZ: Understands that "asking for forgiveness" is the correct cultural behavior, even if the wording is slightly different.
3. The "Magic Translator" (Post-Processing)
Persian is tricky because it has many variations. To fix this, the team built a digital translator that acts like a super-smart editor before grading the AI.
- The Analogy: Imagine you are grading a student's essay, but the student writes "56" while the answer key says "۵۶" (Persian numbers), or uses a fancy word for "and" while the key uses a simple one.
- What TARAZ does: Before grading, it normalizes everything. It turns all numbers into the same format, removes unnecessary accents, and splits long sentences into logical chunks. It's like the teacher saying, "Okay, I see you meant '56' and 'and', let's ignore the font differences and look at the meaning."
4. The Three "Cultural Gymnasiums" (Datasets)
To test the AI thoroughly, they used three different types of "gyms" (datasets):
- BLEnD (Everyday Knowledge): Like a trivia night. "What's a common snack for kids?" (Fruit, sandwiches, etc.).
- PerCul-SAQ (Story Time): The AI reads a short story about a cultural situation and has to guess what happens next or what object is being used. It's like reading a mystery novel and having to solve the cultural clue.
- ISN-SAQ (Social Rules): This tests "unwritten rules." "What is the polite thing to do in a crowded bazaar?" (Answer: Bargain). This is the hardest test because it requires understanding social nuance, not just facts.
5. The Results: Who Passed the Test?
The researchers tested 15 different AI models (some famous ones like GPT-5 and Claude, and some smaller, Persian-specific ones).
- The Big Winners: The giant, closed-source models (like GPT-5 and Claude Opus) performed the best. They are like the "Olympic athletes" of AI—they have seen so much data that they intuitively understand the cultural nuances.
- The Surprise: Some smaller, open-source models (like Gemma-2) did surprisingly well, almost catching up to the giants.
- The Struggle: The models specifically "fine-tuned" just for Persian actually did worse than the big general models.
- Why? It's like a student who memorized a dictionary but never actually talked to people. They knew the words but didn't understand the culture. The researchers suggest these models need more real-world cultural training, not just language translation.
6. The "Human vs. Robot" Judge
Finally, they asked: "Who is a better grader: a human or an AI?"
- They used a new AI grader (called Maux) that uses the "Magic Translator" mentioned earlier.
- Result: The new AI grader agreed with human judges 95% of the time.
- The Takeaway: The old "exact match" method was like a robot with a rigid rulebook. The new TARAZ method is like a human teacher who understands context, slang, and different ways of saying the same thing.
Summary
TARAZ is a new, fairer way to test if AI understands Persian culture. It stops the AI from cheating by memorizing exact answers and forces it to actually understand the meaning behind the words. It found that while the biggest AI models are getting good at this, the specialized Persian models still have a lot of learning to do to truly grasp the human side of the language.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.