← Latest papers
💬 NLP

Evaluating Monolingual and Multilingual Large Language Models for Greek Question Answering: The DemosQA Benchmark

This paper introduces the DemosQA benchmark, a novel Greek question-answering dataset derived from social media, and presents an extensive evaluation of 11 monolingual and multilingual Large Language Models to assess their effectiveness in capturing Greek social and cultural nuances.

Original authors: Charalampos Mastrokostas, Nikolaos Giarelis, Nikos Karacapilidis

Published 2026-02-20
📖 4 min read☕ Coffee break read

Original authors: Charalampos Mastrokostas, Nikolaos Giarelis, Nikos Karacapilidis

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a giant, super-smart robot brain (a Large Language Model, or LLM) that has read almost everything on the internet. This brain is amazing at answering questions in English, but when you ask it about Greek culture, history, or daily life, it often stumbles. It's like asking a world-famous chef who only knows how to cook Italian pasta to make a traditional Greek moussaka; they might get the ingredients right, but they'll miss the soul of the dish.

This paper is about fixing that problem for the Greek language. Here is the story of what the researchers did, explained simply:

1. The Problem: The "One-Size-Fits-All" Trap

Most of these super-smart robots are trained mostly on English data. When they try to learn other languages (like Greek), they often just "translate" what they know from English.

  • The Analogy: Imagine trying to learn Greek by only reading English books about Greece. You might know the facts, but you won't understand the jokes, the slang, or the deep cultural feelings. The robots were missing the "Greek soul."

2. The Solution: A New "Greek Brain" and a New Test

The researchers wanted to see if they could build or find a robot that truly understands Greek. To do this, they did three main things:

A. Building a New Test: "DemosQA"

They needed a way to test the robots fairly. Existing tests were like standardized school exams—good, but they didn't capture how real people actually talk.

  • The Analogy: Instead of giving the robots a textbook quiz, the researchers went to the digital town square (Reddit's Greek community) and collected real questions regular people ask each other, along with the best answers the community voted on.
  • The Result: They created DemosQA, a new dataset that feels like a lively Greek conversation rather than a stiff exam. It covers everything from politics to daily life, capturing the true "zeitgeist" (the mood) of Greece.

B. The "Budget-Friendly" Lab

Usually, testing these giant robots requires super-expensive, massive computers (like a data center). The researchers wanted to make this accessible to everyone.

  • The Analogy: They figured out how to shrink the giant robots down so they could fit into a standard laptop without losing much of their brainpower. It's like taking a massive, fuel-hungry truck and converting it into a fuel-efficient hybrid car that can still drive up a mountain. They used a trick called "4-bit quantization" to make this possible.

C. The Big Showdown: 11 Robots vs. 6 Tests

They lined up 11 different robot brains:

  • The "Greek-Native" Robots: Models specifically trained on Greek text (like Llama Krikri and Meltemi).
  • The "Multilingual" Robots: Models that know many languages but aren't Greek specialists (like Gemma and Aya).
  • The "Rich" Robot: A proprietary model called GPT-4o mini (the expensive, closed-source one).

They tested them on 6 different Greek datasets (including their new DemosQA) using three different ways of asking questions (like giving a direct order, pretending to be a specific character, or asking the robot to "think step-by-step").

3. The Results: Who Won?

The results were surprising and encouraging:

  • The "Rich" Robot (GPT-4o mini) was still the overall champion, especially on hard, specialized topics like Greek law, medicine, and civil service exams. It's the "Olympic Gold Medalist."
  • The "Greek-Native" Robots (specifically Llama Krikri) did incredibly well. In the new, real-world "town square" test (DemosQA), they actually tied with or beat the expensive GPT-4o mini!
  • The Takeaway: You don't always need the most expensive, closed-source robot. A smaller, open-source robot that was specifically trained on Greek culture can be just as smart, if not smarter, for everyday Greek questions.

4. The "Secret Sauce" (Prompting)

The researchers also found that how you ask the question matters:

  • GPT-4o mini works best with simple, direct instructions.
  • The Open-Source Greek models needed a little more help. If you told them, "You are a wise Greek historian," they performed much better. It's like giving a student a specific role to play; they suddenly know exactly what to do.

Summary

This paper is a victory for the Greek language in the world of AI.

  1. They built a new, realistic test (DemosQA) based on real people's questions.
  2. They proved that you can run these tests on regular computers, not just supercomputers.
  3. They showed that open-source models trained specifically for Greek are catching up to the expensive giants, especially when it comes to understanding culture and real-life conversations.

In short: The Greek language is no longer an afterthought for AI. With the right tools and data, we can build robots that truly understand the Greek way of life.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →