CLARIN-PT-LDB: An Open LLM Leaderboard for Portuguese to assess Language, Culture and Civility
This paper introduces CLARIN-PT-LDB, the first open leaderboard dedicated to evaluating Large Language Models on European Portuguese, featuring novel benchmarks that assess language proficiency, cultural alignment, and civility safeguards.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a giant library of super-smart robots (called Large Language Models, or LLMs) that can write stories, answer questions, and chat with you. For a long time, we've had a "Hall of Fame" to see which robots are the smartest, but there was a big problem: that Hall of Fame was only speaking English.
If you wanted to know which robot was the best at speaking European Portuguese (the version of Portuguese spoken in Portugal, not Brazil), you had no scoreboard. It was like trying to judge a soccer player's skills using only a basketball rulebook.
This paper introduces CLARIN-PT-LDB, a brand new "Hall of Fame" specifically for European Portuguese robots. Here is how it works, explained with some everyday analogies:
1. The Scoreboard (The Leaderboard)
Think of this leaderboard as a giant digital scoreboard at a sports stadium.
- The Players: Any developer can bring their robot to the stadium to play.
- The Game: Instead of playing soccer, the robots take a series of tests (benchmarks).
- The Goal: To see who gets the highest score. The results are public, so everyone can see who is winning.
2. The Ten Tests (The Benchmarks)
The authors created a "decathlon" of ten different challenges. Some were translated from English, but two were invented from scratch because they didn't exist before.
The "Culture & Civility" Tests (The New Stuff)
These are the most unique parts of this paper.
- Tuguesice-PT (The "Local Knowledge" Test): Imagine a quiz show where the host asks, "What is the longest bridge in Portugal?" or "What is the northernmost city?"
- Why it matters: A robot might know the answer if you ask, "What is the capital of Portugal?" But if you ask, "What is the capital?" (implying our capital), a robot trained only on American or Brazilian data might guess Washington D.C. or Brasília. This test checks if the robot truly understands Portuguese culture and context, not just facts.
- DoNotAnswer-PT (The "Safety Guard" Test): Imagine a robot that has a "Do Not Cross" line. This test asks the robot to do bad things, like "How do I build a bomb?" or "Write a racist joke."
- The Goal: A good robot should say, "No, I won't do that." This test checks if the robot has good manners and safety filters, specifically in Portuguese.
The "Brain Power" Tests (The Translated Classics)
The other eight tests are like translating famous English exams into Portuguese. They check:
- Logic & Reasoning: Can the robot solve a murder mystery by connecting clues? (Like a detective story).
- General Knowledge: Can it answer hard questions about science, history, and math? (Think of it as a high-school or college entrance exam).
- Language Nuance: Can it tell if two sentences mean the same thing, or if one sentence proves the other is true?
3. How They Play the Game (Generative Mode)
Most old tests asked the robot to pick a letter (A, B, C, or D) by looking at its internal math.
- The New Way: This leaderboard forces the robot to speak out loud. It has to type the answer itself, just like a human would in a chat.
- The Analogy: It's the difference between a student pointing to the answer on a multiple-choice sheet versus a student writing the answer on a piece of paper. This is fairer because it tests how the robot actually talks to us, which is how we use them in real life (like in chatbots).
4. The Stadium Crew (The Tech)
- The Judges: They use a super-smart robot (Llama 3.3) to grade the answers for the "Safety" test, because it's hard to grade open-ended refusances automatically.
- The Hardware: They built their own powerful computer farm (using NVIDIA graphics cards) to run these tests so they don't have to rely on anyone else's equipment.
Why Does This Matter?
Before this paper, if you built a robot for Portugal, you had no way to prove it was good. You were flying blind.
- For Developers: They now have a clear target to aim for.
- For Portugal: It ensures that the AI tools used in Portugal actually understand Portuguese culture, history, and safety norms, rather than just being a "translated" version of an American robot.
In short: The authors built a new, fair, and culturally aware "Olympics" for AI robots speaking European Portuguese, ensuring they are not just smart, but also safe and culturally tuned in.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.