OMGEval: An Open Multilingual Generative Evaluation Benchmark for Large Language Models
This paper introduces OMGEval, the first open-source multilingual generative evaluation benchmark featuring 804 rigorously verified, culturally localized open-ended questions across five languages to assess and improve the multilingual capabilities of large language models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have built a brilliant, all-knowing librarian who can speak every language in the world. You want to test if this librarian is truly smart in all cultures, or if they are just a "local expert" who only understands the culture of the United States.
This paper introduces OMGEval, which is essentially a multicultural "driver's license" test for these AI librarians (Large Language Models).
Here is the breakdown of what the authors did, using simple analogies:
1. The Problem: The "American TV" Bias
Most current tests for AI are like watching a TV show that only airs in English and talks about American holidays, food, and history.
- The Issue: If you ask an AI, "What is the traditional food for Thanksgiving?" it will correctly say "Turkey." But if you ask a Chinese speaker, "What is the traditional food for the Dragon Boat Festival?" an AI trained only on American data might get confused or give a wrong answer because it doesn't understand the local culture.
- The Paper's Claim: Most existing tests are too focused on English. They don't check if the AI understands that in Russia, Easter involves specific cakes, or in Spain, there are specific treats for All Saints' Day.
2. The Solution: OMGEval (The "Local Tour Guide" Test)
The authors created a new test called OMGEval. Instead of just translating English questions into other languages (which is like putting English subtitles on a Russian movie), they rewrote the questions to fit the local culture.
- The Analogy: Imagine a test question about "Summer Vacation."
- The Old Way (Translation): "Where do Americans go for summer vacation?" (Answer: Hawaii).
- The OMGEval Way (Localization): "Where do Chinese people go for summer vacation?" (Answer: Sanya).
- Why it matters: Both questions ask about "summer vacation spots," but the cultural context is different. OMGEval ensures the AI knows the local version of the story, not just the American one.
3. How They Built It
The team didn't just use a robot to translate things; they used a team of human experts (linguists) to act as cultural editors.
- The Process: They took 804 open-ended questions (like riddles, facts, or creative writing prompts) and adapted them for 5 languages: Chinese, Russian, French, Spanish, and Arabic.
- The "Human Touch": If a question mentioned a specific American celebrity, they swapped it for a famous local figure. If it mentioned a specific US holiday, they swapped it for a local festival. They did this to make sure the AI was being tested on its ability to understand that specific culture, not just its ability to translate words.
4. How They Graded the AI
Since humans can't read thousands of answers in five languages instantly, they used GPT-4 (a very advanced AI) as the referee.
- The Game: They asked the AI models to answer the questions. Then, they asked GPT-4 to look at two answers side-by-side and decide which one was better.
- The Check: They also had real humans grade a small sample to make sure the "AI Referee" was being fair. The results showed the AI Referee was very accurate (about 90% agreement with humans).
5. The Results: Who Passed the Test?
They tested several different AI models, including big commercial ones (like GPT-4) and open-source ones (free models anyone can download).
- The Winner: GPT-4 was the only model that consistently scored above 50% (meaning it beat the baseline more than half the time). It was the "star student."
- The Struggle: The open-source models (like Guanaco, Phoenix, and others) struggled significantly.
- The Gap: While GPT-4 could handle cultural nuances well, the open-source models often made factual errors or missed the cultural context. For example, when asked about a famous Chinese director's first movie, some open-source models gave the wrong year or the wrong movie title, while GPT-4 got it right.
- The Takeaway: There is a huge gap between the top-tier commercial models and the open-source models when it comes to understanding different cultures. The open-source models are like students who studied the textbook but failed to understand the local dialect.
Summary
OMGEval is a new, fairer test that checks if AI models understand the world's diverse cultures, not just the American version of the world. The paper shows that while the best AI (GPT-4) is quite good at this, many other models are still struggling to understand the local "flavor" of different languages and cultures. The authors hope this test will help the community build better, more culturally aware AI in the future.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.