mmPISA-bench: Do LLMs Reason Equally Well Across 43 Languages?
This paper introduces mmPISA-bench, a multilingual reasoning benchmark derived from PISA across 43 languages, demonstrating that modern LLMs achieve human-comparable reasoning performance with minimal degradation from machine translation, though revealing significant variations in accuracy and inference costs across different languages.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart student who can speak almost any language. You want to know: Is this student equally smart when taking a test in French as they are when taking it in English? Or, if you give them a test written in a language they don't know natively (but translated by a robot), do they still get the right answers?
This paper, called mmPISA-bench, is like a massive, multilingual report card for two of the world's most advanced AI "students" (specifically, models from OpenAI and Anthropic). Here is the breakdown of what they found, using some everyday analogies.
1. The Test: A Global Math and Reading Exam
The researchers didn't just make up random questions. They grabbed 25 real questions from the PISA exam—the famous international test that 15-year-old students around the world take to prove they can solve math problems and understand reading passages.
- The Setup: They took these 25 questions and translated them into 43 different languages.
- The Twist: For every language, they had two versions:
- The "Official" Version: Translated by human experts (like a professional translator).
- The "Robot" Version: Translated by Google Translate (machine translation).
- The Goal: See if the AI could solve these logic puzzles in all 43 languages, and if it mattered whether the translation was done by a human or a robot.
2. The Results: Are the AI Students Equally Smart Everywhere?
The Good News: Yes, mostly.
The AI models performed very well across all 43 languages. In fact, their accuracy was often comparable to what you'd expect from a bright 15-year-old human student. They didn't suddenly "forget" how to do math just because the question was in Icelandic or Thai.
The Catch: It wasn't perfectly equal.
- The "Dialect" Effect: Just like a human might be slightly more comfortable reading a story in their native tongue than in a second language, the AI was slightly more accurate in some languages than others.
- The "Robot" Translation Myth: You might think, "If the question is translated by a robot, the AI will get confused." The paper says no. The AI got the right answers just as often with the "Robot" translations as with the "Human" ones. This is a big deal because it means we don't always need expensive human translators to test AI in new languages; a good machine translator works fine.
3. The Hidden Cost: The "Thinking" Price Tag
This is where the paper gets really interesting. It's not just about getting the answer right; it's about how much it costs to get there.
Imagine two students taking a test:
- Student A (Claude): Writes a very long, detailed explanation for every answer.
- Student B (GPT): Writes shorter explanations.
The researchers found that for some languages, the AI had to "think" much harder (write more text) to get the right answer, which made the test more expensive to run.
- The "Expensive & Clumsy" Languages: For certain languages (like Thai for one model, or Greek for the other), the AI had to generate a huge amount of text to solve a simple problem. It was like paying for a luxury taxi ride just to go to the corner store.
- The Result: In these "expensive" languages, the AI was actually less accurate and more costly at the same time. It's a double whammy: you pay more, and you get a worse result.
4. The Secret Language Switch
The researchers also peeked under the hood to see how the AI was thinking. They found something surprising:
- Sometimes, when the AI was asked a question in Kazakh, it didn't just think in Kazakh or English. It actually started doing its math and logic in Russian before giving the answer back in Kazakh.
- It's like a person who speaks Spanish and English, but when they get a hard math problem, they instinctively switch to thinking in French because that's where their "math brain" lives. The AI was doing something similar, using a "pivot language" to solve the problem.
Summary: What Does This Mean?
- AI is getting good at languages: Modern AI can reason through logic puzzles in 43 different languages almost as well as a human teenager.
- Machine translation is okay: You don't need perfect human translations to test AI; robot translations work just fine for this.
- Language matters for cost: Some languages are "harder" for AI to process. They make the AI talk more (which costs money) and sometimes make it make more mistakes.
- We need to look deeper: To truly understand how good an AI is, we can't just look at the final score. We have to look at how much "effort" (and money) it took to get that score, and what languages it was secretly thinking in.
In short, the AI is a polyglot genius, but it still has some "dialects" where it stumbles a bit and spends a lot more of its allowance to get the job done.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.