← Latest papers
💬 NLP

GreekMMLU: A Native-Sourced Multitask Benchmark for Evaluating Language Models in Greek

This paper introduces GreekMMLU, a comprehensive, native-sourced benchmark comprising over 21,000 multiple-choice questions across 45 subjects, designed to address the lack of authentic evaluation tools for Greek language models and revealing significant performance gaps between frontier, open-weight, and Greek-adapted models.

Original authors: Yang Zhang, Mersin Konomi, Christos Xypolopoulos, Konstantinos Divriotis, Konstantinos Skianis, Giannis Nikolentzos, Giorgos Stamou, Guokan Shang, Michalis Vazirgiannis

Published 2026-03-31
📖 4 min read☕ Coffee break read

Original authors: Yang Zhang, Mersin Konomi, Christos Xypolopoulos, Konstantinos Divriotis, Konstantinos Skianis, Giannis Nikolentzos, Giorgos Stamou, Guokan Shang, Michalis Vazirgiannis

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a giant library of knowledge, and inside it are thousands of books written in English. Now, imagine you want to test how smart a new AI robot is. You give it a quiz based on those English books. If the robot gets a high score, you think, "Wow, this robot is a genius!"

But here's the problem: What if you want to test that same robot on Greek culture, history, and language?

If you just take the English quiz and use a machine translator to turn it into Greek, the robot might still get a high score. But it's not because it understands Greek; it's just because it's good at guessing based on the English version it already knows. It's like asking someone to recite a poem in a language they don't speak, but they memorized the English translation and just read the words backwards. They sound like they know the language, but they don't really get the meaning, the jokes, or the cultural heart of it.

This is exactly the problem the authors of this paper wanted to fix.

The Problem: The "Translated" Trap

For a long time, to test AI on Greek, researchers just took English tests and translated them. This is like trying to judge a chef's ability to cook authentic Greek food by giving them a recipe translated from French. The ingredients might be there, but the flavor is wrong. The AI misses the nuance, the specific way Greeks think about their history, and the unique structure of their language.

The Solution: GreekMMLU (The "Native" Quiz)

The team created GreekMMLU, which is like building a brand-new, authentic Greek test from scratch.

  • Where did the questions come from? Instead of translating English, they went straight to the source: real Greek school exams, university tests, professional licensing exams (like for doctors or lawyers), and even driving tests.
  • What's inside? It's a massive collection of 21,805 questions covering 45 different subjects.
    • STEM: Math, Physics, Biology.
    • Humanities: Greek Mythology, History, Literature.
    • Social Sciences: Economics, Law, Politics.
    • Greek-Specific: Things that only a Greek speaker would know, like the rules of Greek driving or the details of Greek Orthodox traditions.
  • The "Secret Sauce": They split the questions into two piles. One pile is public (so anyone can test their AI), and the other is a "secret" pile kept in a vault. This prevents AI companies from accidentally memorizing the test answers before the test happens (a problem called "data contamination").

The Big Experiment: Who Passed the Test?

The authors took over 80 different AI models (both free, open-source ones and expensive, closed-source ones) and gave them this Greek quiz. Here is what they found, using some simple analogies:

  1. The "Big Brains" vs. The "Small Brains":

    • The biggest, most powerful AI models (like the ones from Google and OpenAI) did very well. They are like students who have read every book in the library.
    • The smaller, open-source models struggled. Some of them scored barely better than random guessing. It's like giving a high school physics test to someone who hasn't finished middle school yet.
  2. The "Specialist" vs. The "Generalist":

    • Some AIs are "Generalists"—they know a little bit about everything in many languages.
    • Some are "Specialists"—they were specifically trained to speak and understand Greek better.
    • The Result: The Specialists won! Models that were specifically tuned for Greek (like Llama-Krikri or Meltemi) performed much better than general multilingual models of the same size. It's like hiring a local tour guide (Specialist) versus a guide who just learned the language from a phrasebook (Generalist). The local guide knows the shortcuts and the hidden gems.
  3. The "Instruction" Boost:

    • They found that if you "teach" the AI how to take the test (by giving it examples of how to answer), it gets smarter. This is called "instruction tuning." It's like giving a student a practice exam before the real one; they do much better because they understand the format.

Why Does This Matter?

This paper is a wake-up call. It shows that just because an AI speaks Greek, doesn't mean it understands Greek.

  • For Researchers: It gives them a fair, honest ruler to measure how good an AI really is at Greek.
  • For the Future: It encourages companies to stop just translating English tests and start building real, native-language tests for every culture. If we want AI to be truly helpful to Greeks, it needs to understand Greek culture, not just Greek words.

In short: The authors built a real, authentic Greek exam to stop AI from "faking it." They found that while the biggest AIs are smart, the ones specifically trained on Greek culture are the true champions of the language.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →