Truth Knows No Language: Evaluating Truthfulness Beyond English
This paper introduces a professionally translated extension of the TruthfulQA benchmark for Basque, Catalan, Galician, and Spanish to evaluate 12 state-of-the-art LLMs, revealing that while performance varies by resource level, truthfulness discrepancies across languages are smaller than expected and that machine translation offers a scalable alternative for cross-lingual truthfulness assessment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, well-read librarian who speaks perfect English. You ask them tricky questions like, "Is it illegal to burn a flag in the US?" or "Do chameleons change colors to hide?" The librarian knows the answers and can explain why common myths are wrong. This librarian is like the Large Language Models (LLMs) we use today.
But what happens if you ask that same librarian these questions in Basque, Catalan, Galician, or Spanish? Do they still know the truth, or do they start making things up because they haven't read as many books in those languages?
This paper is like a massive, multilingual "truth test" designed to find out. Here is the story of what they did and what they found, explained simply.
The Big Problem: The English Bias
For a long time, we've only tested these AI librarians in English. It's like testing a chef only on how well they cook steak, then assuming they can cook a perfect sushi roll just because they are a "good chef." The researchers wanted to know: Does the AI tell the truth equally well in every language, or does it get confused when the language changes?
They took a famous English test called TruthfulQA (which is full of common myths and misconceptions) and had professional human translators turn it into Basque, Catalan, Galician, and Spanish. This is important because they didn't just use a robot to translate it; they used humans to make sure the cultural nuances stayed intact.
The Experiment: 12 Librarians, 5 Languages
They put 12 different AI models (the "librarians") through this test. They asked them questions in all five languages and used three different ways to grade them:
- Human Grading: Real people read the answers and decided if they were true and helpful.
- Multiple Choice: A computer checked if the AI picked the right pre-written answer (like a bubble sheet).
- The "AI Judge": They used a super-smart AI to grade the other AIs, acting like a referee.
The Surprising Findings
1. The "Resource Gap" is Real, but Smaller Than Expected
Think of language resources like food. English is a five-star banquet; Basque is a small, simple snack. The researchers found that the AIs performed best at the banquet (English) and worst at the snack table (Basque).
- The Twist: The difference wasn't as huge as everyone feared. The AIs didn't completely lose their minds in the smaller languages; they just got a little less precise.
2. The "Silent Librarian" Trick
Here is a funny quirk they found. Some of the older, "base" AI models (the ones not trained to chat) would often answer tricky questions with, "I have no comment."
- The Trap: In the test rules, saying "I don't know" counts as being truthful (because it's better than lying). So, these models got high scores just by staying silent.
- The Reality: When the researchers looked closer, they saw these "silent" models weren't actually being helpful; they were just refusing to answer. The newer, "instruction-tuned" models (the chatty ones) actually gave better, more honest answers, even if they were a bit more complex.
3. The "AI Judge" is the Best Referee
The researchers tried to see which grading method was best.
- Multiple Choice was like a rigid robot: it only checked if the answer matched a list. It missed the nuance.
- The AI Judge was like a smart human referee. It correlated much better with what real humans thought was "truthful." Even if the Judge was trained in English, it could spot truthfulness in Spanish or Basque better than the multiple-choice method.
4. Bigger Isn't Always Better (But Usually Is)
In some past studies, bigger AI models didn't always do better. But here, the researchers found that the bigger models (70 billion parameters) were consistently more truthful than the smaller ones (8 billion), especially in the non-English languages. It seems that having a bigger "brain" helps the AI navigate the tricky waters of different languages.
5. Universal vs. Local Knowledge
The test had two types of questions:
- Universal: "Why do chameleons change colors?" (True everywhere, always).
- Local/Time-Sensitive: "Is burning a flag illegal in the US?" (Specific to a place and time).
The AIs were great at the universal questions (almost 90% accuracy). But they struggled more with the local, time-sensitive ones. This suggests that if we only test AIs on "eternal truths," we aren't really testing how well they handle the messy, changing real world.
6. Robots Can Translate the Test (Maybe)
Finally, they wondered: "Do we need expensive human translators to make these tests for new languages?" They tried using a top-tier AI to translate the test instead of humans.
- The Result: The AI-translated test gave almost the exact same results as the human-translated one. This means we might be able to quickly expand these truth tests to hundreds of languages using robots, saving a lot of money and time.
The Bottom Line
The paper concludes that while AI is still a bit "Anglocentric" (better at English), it is surprisingly capable of telling the truth in other languages, provided we use the right tools to measure it.
- Don't trust the "Silent Librarian": Just because an AI says "I don't know" doesn't mean it's being honest; it might just be lazy.
- Use a Smart Judge: Let a smart AI grade the answers; it's better than a simple checklist.
- Watch out for Local Context: AIs are great at facts that never change, but they need more practice with facts that depend on where and when you ask them.
The researchers made all their data, models, and code public, so anyone can try to teach their own AI librarian to tell the truth in any language.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.