← Latest papers
💬 NLP

MGSM-Pro: A Simple Strategy for Robust Multilingual Mathematical Reasoning Evaluation

This paper introduces MGSM-Pro, a multilingual extension of the GSM-Symbolic benchmark that evaluates mathematical reasoning robustness across nine languages by generating multiple digit-varying instantiations, revealing significant performance drops in low-resource languages and varying model resilience to such perturbations.

Original authors: Tianyi Xu, Kosei Uemura, Alfred Malengo Kondoro, Tadesse Destaw Belay, Catherine Nana Nyaah Essuman, Ifeoma Okoh, Ganiyat Afolabi, Ayodele Awokoya, David Ifeoluwa Adelani

Published 2026-04-29
📖 5 min read🧠 Deep dive

Original authors: Tianyi Xu, Kosei Uemura, Alfred Malengo Kondoro, Tadesse Destaw Belay, Catherine Nana Nyaah Essuman, Ifeoma Okoh, Ganiyat Afolabi, Ayodele Awokoya, David Ifeoluwa Adelani

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a teacher giving a math test to a class of students from different countries. You want to see who is truly good at math and who is just memorizing the answers from a practice book.

This paper introduces a new, tougher way to test Large Language Models (AI) on math problems, called MGSM-Pro. Here is the story of what they found, explained simply.

The Problem: The "Practice Book" Trap

For a long time, AI models have been getting better at solving math word problems. But researchers noticed something suspicious: the AI might not actually be "thinking" through the math. Instead, it might be memorizing specific details from the test questions it saw during training.

Think of it like a student who memorizes the answer to a specific word problem: "If John has 5 apples and buys 3 more..." The student knows the answer is 8. But if you change the question to "If Sarah has 5 oranges and buys 3 more...", the student who only memorized the first one might get confused.

Recently, researchers found that even in English, if you change the names or numbers in a math problem slightly, many AI models fail. They realized that the old tests were too easy because they didn't check if the AI could handle these small changes.

The Solution: MGSM-Pro (The "Shuffle" Test)

The authors created a new dataset called MGSM-Pro. They took 250 math problems and turned them into 1,240 problems (5 versions of each).

They did this by shuffling the deck:

  1. Changing Names: Swapping "John" for "Zainabu" or "Carla."
  2. Changing Numbers: Swapping "5 apples" for "7 apples."
  3. Adding Distractions: Adding a sentence that sounds important but is actually nonsense (like "The sky is blue, so John bought 5 apples").

They did this in nine different languages, ranging from widely spoken ones like English and Chinese to languages with fewer digital resources like Swahili, Yoruba, and Igbo.

The Big Discovery: The "Low-Resource" Struggle

When they ran the tests, they found a massive gap between "High-Resource" languages (like English) and "Low-Resource" languages (like Twi or Igbo).

  • The High-Resource Students: When you changed the names or numbers in English or Chinese, the AI models dropped a little bit in score, but they stayed mostly on track. They were like a student who actually understood the math concept.
  • The Low-Resource Students: When they did the same shuffling for languages like Twi or Igbo, the AI models crashed. Their scores plummeted. It was as if the student suddenly forgot how to count because the teacher used a different name for the fruit.

The Analogy: Imagine a student who is great at math in English but terrible at math in their native language. When the teacher asks a question in English with a twist, the student solves it. But when the teacher asks the exact same twisted question in the student's native language, the student freezes. The paper found that AI models are doing the same thing: they are much less "robust" (reliable) in languages they haven't seen as much data for.

Who Passed and Who Failed?

The paper tested many different AI models (some made by big tech companies, some open-source).

  • The Robust Ones: Newer, smarter models like Gemini 3.0 Pro and GPT-OSS 120B handled the changes well. They didn't panic when the numbers or names changed.
  • The Fragile Ones: Older models or smaller models (like Gemma 3 4B or Llama 3) fell apart. They couldn't handle the "twist."
  • The Size Myth: You might think a bigger brain (more parameters) always means a smarter student. The paper found this isn't true. Sometimes a huge model failed miserably, while a slightly smaller one succeeded. It's not just about size; it's about how the model was trained.

The Lesson: Don't Just Test Once

The biggest takeaway is about how we should test AI in the future.

Currently, many leaderboards (rankings of the best AI) test a model on just one version of a problem. The authors say this is like giving a student a single practice question and calling them a genius if they get it right.

Their Recommendation: To truly know if an AI is good at math, you must test it on at least five different versions of the same problem (changing the names and numbers). If the AI gets all five right, then it's truly smart. If it only gets the first one right, it's just memorizing.

Summary

The paper argues that we need to stop treating AI like it's perfect just because it passes the easy tests. By shuffling the names and numbers in math problems across different languages, they showed that:

  1. AI is much weaker in languages with less digital data.
  2. Many current rankings are misleading because they don't test for this "shuffling."
  3. To get a real picture of an AI's math skills, we need to test it on multiple variations of the same problem, not just one.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →