M3Kang: Evaluating Multilingual Multimodal Mathematical Reasoning in Vision-Language Models
This paper introduces M3Kang, the first massively multilingual, multimodal mathematical reasoning dataset derived from the Kangaroo Math Competition, which benchmarks state-of-the-art vision-language models against human performance and reveals their current limitations in basic math and diagram-based reasoning across 108 languages.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a giant, global math contest called the "Kangaroo Math Competition." Every year, over six million kids under 18 from more than 90 countries take part. They solve tricky math puzzles, some of which are just words, but many require looking at diagrams, shapes, and pictures to figure out the answer.
Now, imagine a team of researchers from Qualcomm AI Research decided to turn this contest into a massive test for Artificial Intelligence. They built a new tool called M3Kang. Think of M3Kang as a "universal translator and math teacher" rolled into one dataset.
Here is the simple breakdown of what they did and what they found:
1. The Big Project: Building the Ultimate Math Test
The researchers took the original math problems (which were mostly in Catalan and English) and translated them into 108 different languages.
- The Scale: They didn't just translate the words; they kept the pictures and diagrams attached to the questions. This resulted in over 111,000 unique math problems.
- The Goal: They wanted to see if AI models (the "brains" behind chatbots and image generators) could solve math problems in languages other than English, and if they could actually "see" and understand the diagrams, not just read the text.
2. How They Built It (The Assembly Line)
Creating a dataset this big is like building a factory assembly line.
- Step 1: They took the original PDFs of the math contest and turned them into clean images with text.
- Step 2: They used AI to translate the questions into English first, double-checking that the translation didn't break the math logic.
- Step 3: They used a clever "back-translation" trick to check the other 108 languages. Imagine translating a sentence from Spanish to English, and then immediately translating it back to Spanish. If the result looks different from the original, the translation is bad. They used this method to filter out bad translations automatically.
3. The Results: How Did the AI Do?
They tested the smartest AI models available (both free/open ones and expensive/closed ones) against this new test. Here is what they discovered:
- The "English Advantage" is Real: Just like a student who studies only in English might struggle in a French classroom, the AI models were much better at solving math problems in English. As the language became less common on the internet (like Swahili or Maltese), the AI's performance dropped significantly. It's as if the AI forgot how to do math when the language changed.
- Size Matters (But Not for Everything): Bigger AI models generally did better. However, even the biggest, smartest models struggled with basic logic and diagrams.
- The "Blind" Problem: This was the most surprising finding. When humans take the test, they actually find the problems with pictures easier than the ones with just text. But for AI, it's the opposite! The AI models got significantly worse when a picture was involved. It's like a student who is great at reading a word problem but freezes up when they have to look at a graph. The AI seems to have trouble connecting the dots between the text and the image.
- AI vs. Humans: They compared the AI scores against the real scores of 68,000 human students.
- The best AI (Gemini-2.5-Pro) performed like a top-tier student in English, but its performance dropped to "average" or "below average" in low-resource languages.
- Interestingly, the AI didn't get better as the math got harder in the same way humans do. Sometimes the AI aced the hardest problems but failed the easiest ones, suggesting it's "guessing" or using a different kind of reasoning than a human child.
4. Can We Fix It?
The researchers tried a few "training tricks" to help the AI speak other languages better.
- The "Translator" Trick: They found that if they told the AI, "First, translate this question into English in your head, solve it, then give the answer," the AI got much better at solving problems in other languages. It's like giving a student a dictionary right before the test.
- The "Steering" Trick: They tried to nudge the AI's internal thinking process to be more "English-like" even when speaking other languages. This helped a little, but the "Translator" trick was the winner.
The Bottom Line
The paper concludes that while AI is getting incredibly smart, it still has a major blind spot: it struggles to combine reading, seeing, and thinking across different languages.
The researchers made all their data and code public (open-source) so other scientists can use this "Kangaroo Test" to build better, more inclusive AI that doesn't just speak English, but can actually do math in any language, with or without a picture. They hope this helps create AI that is truly global and fair for everyone.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.