Multilingual Large Language Models do not comprehend all natural languages to equal degrees
This study challenges the assumption that English is the optimal language for Large Language Models by demonstrating that, while these models outperform human baselines less across all tested languages, they often achieve higher comprehension accuracy in certain Romance languages than in English due to factors like tokenization, training data composition, and linguistic distance.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart, multi-lingual robot librarian named "LLM." This robot has read almost every book on the internet and can answer questions in dozens of languages. For years, we've assumed that if you ask this robot a question in English, it will be at its absolute peak performance, like a chef cooking their signature dish. We assumed that for other languages, especially those spoken by fewer people, the robot would be a bit clumsy, like a chef trying to cook a dish they've only seen in a blurry photo.
But this new study is like a reality check for that robot. The researchers decided to test the robot's "comprehension" (its ability to truly understand a story and answer a question about it) across 12 different languages, from English and Spanish to Japanese and Arabic. They compared the robot's answers to those of real human readers.
Here is what they found, explained through some simple analogies:
1. The Robot is Good, But Not Human-Perfect
The Analogy: Imagine a student taking a reading comprehension test. The human students get almost 100% of the answers right. The robot, even the most advanced one, gets about 70–90% right.
The Finding: No matter which language you use, the robot still falls short of human understanding. It's not a perfect mind; it's a very good pattern-matcher. However, the gap between the robot and humans isn't the same everywhere. In some languages, the robot is surprisingly close to human levels; in others, it stumbles.
2. The "English is King" Myth is Dead
The Analogy: For a long time, we thought English was the robot's "native tongue" and its home base. We thought if you asked it in English, it would be sharp as a tack, but if you asked it in French or Japanese, it would have to do a mental "translation" first, making it slower and less accurate.
The Finding: The study flipped the script. English was actually in the middle of the pack! It wasn't the best.
- The Surprise Winners: The robot performed best in Spanish and Italian. It was like finding out that a chef who claims to be French is actually a master of Italian cuisine.
- The Losers: The robot struggled the most with Japanese and Greek, even though these are major world languages.
3. Why Did Spanish Win and English Lose?
The Analogy: Think of the robot's brain as a library where books are stored in "chunks" called tokens.
- The Token Problem: Imagine trying to read a book where every word is broken into tiny, weird pieces. In English, the robot often has to chop words into many small pieces (tokens) to understand them. In Spanish and Italian, the words flow together more naturally for the robot's specific way of "chopping" text.
- The Result: Because the robot can process Spanish and Italian more efficiently (fewer "chunks" to manage), it understands the story better. English, ironically, is a bit "messier" for the robot's internal mechanics.
4. The "WEIRD" Bias (The Western Bubble)
The Analogy: Imagine the robot was trained mostly on books written by wealthy, educated people in Western countries (a group researchers call WEIRD: Western, Educated, Industrialized, Rich, Democratic).
The Finding: You might think this means the robot is great at Western languages. But the study shows that just because a language is "Western" (like English or German) doesn't guarantee the robot understands it perfectly.
- The Real Issue: The robot struggles most with languages that use non-Latin scripts (like Japanese, Chinese, or Arabic) and have less data available. It's like the robot has a pair of glasses that are perfectly tuned for Latin letters but are slightly blurry for other scripts. If a language is both "non-Latin" and "low-resource," the robot is essentially trying to read a map in the dark.
5. Stability: The Robot is a Robot, Humans are Humans
The Analogy: Imagine asking the same question to a human and a robot 10 times.
- The Human: Sometimes you're tired, distracted, or having an "off" day, so you might answer differently each time.
- The Robot: If you ask the robot the same question 10 times, it usually gives the exact same answer every time.
The Finding: In terms of consistency (stability), the robots were actually more stable than humans in most languages. Humans get distracted; robots just follow their code. However, this consistency doesn't mean they are right—they can be consistently wrong!
The Big Takeaway
This study is a wake-up call. We can't just assume that because a robot speaks English, it understands the world.
- The "Romance Puzzle": Surprisingly, the robot understands Spanish and Italian better than English.
- The Risk: If we rely on these robots for important things (like medical advice or legal help) in languages like Japanese or Greek, we might get answers that are less reliable than we think.
- The Future: We need to stop assuming English is the "gold standard" and realize that the robot's "brain" is shaped by how it was built and what data it ate, not just by how many people speak a language.
In short: The robot is a brilliant polyglot, but it has a favorite accent (Spanish/Italian) and a blind spot (non-Latin scripts). It's not the perfect human substitute we hoped for, and it definitely isn't an English-only genius.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.