GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models
This paper introduces GSM-Plus-BN, a novel perturbation-based benchmark for Bengali mathematical reasoning derived from the English GSM-Plus dataset, to evaluate the robustness and reasoning capabilities of six large language models while highlighting the significant performance gaps and challenges inherent in low-resource language contexts.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to solve math problems. You wouldn't just ask it to calculate ; you'd want to know if it truly understands why the answer is 4, or if it's just memorizing that "2 plus 2" always equals 4. This is the big question in the world of Artificial Intelligence: Do Large Language Models (LLMs)—the super-smart computer brains behind chatbots—actually "think," or are they just really good at guessing the next word in a sentence?
To find out, scientists have been testing these robots with math word problems. But here's the catch: if you ask a robot the exact same question a thousand times, it might just memorize the answer. To see if it's truly smart, you have to trick it. You change the numbers, swap the words, or add confusing extra details to see if the robot still gets the right answer. This is called "perturbation." It's like asking a student, "If I have 3 apples and eat 1, how many are left?" and then, a second later, asking, "If I have 300 apples and eat 100, how many are left?" A smart student knows the math is the same; a robot that just memorized the first answer might get confused.
Now, imagine doing all this testing, but not in English, the language most robots are trained on, but in Bengali, a language spoken by over 230 million people. Until now, there was almost no way to test if these AI brains could handle math tricks in Bengali. This paper steps in to fill that huge gap, asking: Can these AI models really reason in Bengali, or do they just stumble when the words change?
The researchers behind this study, a team from universities in Bangladesh, decided to build a giant, tricky playground for AI called GSM-PLUS-BN. Think of it as a "Math Obstacle Course" specifically designed for the Bengali language. They started with a set of 1,000 standard math problems (the "seed" questions) and then used a clever recipe to create 8,000 new, trickier versions of them.
How did they make these tricks? They used eight different types of "distortions," like a magician changing a card in a deck:
- Changing the Numbers: Swapping a "3" for a "4" or turning a small number into a huge one (like 2 becoming 2,000).
- Changing the Math: Adding a new step to the problem or flipping the question around (instead of asking "What is the total?", asking "If the total is X, what was the starting number?").
- Changing the Words: Rewriting the sentence so it means the same thing but sounds totally different.
- The "Red Herring": Adding a sentence that sounds important but is actually useless, like mentioning the price of a shirt when the question is only about how many shirts you need.
- The "Missing Piece": Removing a crucial piece of information to see if the robot admits it doesn't know the answer or just guesses anyway.
They tested six different AI models on this course. Some were small and fast (like a smart calculator), while others were massive and complex (like a super-computer brain). They asked the robots to solve the problems in two ways: first, just by giving a direct answer, and second, by "thinking step-by-step" (a technique called Chain-of-Thought, where the robot has to show its work).
What did they find?
The results were a mix of "Wow!" and "Uh-oh."
First, the "thinking step-by-step" trick worked wonders for some robots but backfired for others. The Qwen3-32B model was like a student who was terrible at math until the teacher said, "Show your work!" Suddenly, its score jumped from a dismal 13.98% to a solid 76.25%. It seems this model had the potential to be smart, but it needed a little nudge to unlock it.
However, the biggest robots didn't always get the biggest boosts. The GPT-OSS-20B (the 20-billion-parameter model) was actually the best at solving the standard questions without any extra help, getting 96.08% right. But when forced to "think step-by-step," it actually got slightly worse, dropping to 87.80%. Meanwhile, the massive GPT-OSS-120B giant started with a strong 88.03% on standard questions and also saw a slight dip when asked to show its work. It seems these giants were so good at guessing the answer directly that the extra instructions sometimes confused them.
The most surprising discovery was how hard the robots found the "Critical Thinking" traps. When the researchers removed a key piece of information and asked, "Can you solve this?", almost all the models failed miserably. Even the best ones only got about 33% right. This suggests that while these AI models are great at following patterns, they still struggle to realize when a problem is unsolvable. They tend to guess anyway rather than saying, "I don't have enough info."
Another big finding was that math is still hard for AI, no matter the language. Even the best models struggled with problems that required changing how numbers were written (like turning whole numbers into fractions) or adding extra math steps. The paper suggests that these models might be relying too much on recognizing patterns in the text rather than doing the actual math in their "heads."
Finally, the team found that the AI models were much more fragile in Bengali than they are in English. When the questions were changed slightly, the robots' scores dropped significantly. This tells us that while AI is getting smarter, it still has a long way to go before it can truly "understand" math in languages like Bengali the way a human does.
In short, this paper didn't just build a new test; it showed us that while AI can be surprisingly good at math, it's still easily confused by tricks, especially in languages it hasn't mastered as well as English. The "thinking step-by-step" trick helps some, but it's not a magic wand that fixes everything. The robots are getting better, but they aren't quite ready to be the ultimate math geniuses just yet.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.