← Latest papers
💬 NLP

ArithmAttack: Evaluating Robustness of LLMs to Noisy Context in Math Problem Solving

The paper introduces ArithmAttack, a method for evaluating the robustness of large language models to punctuation noise in math problem-solving, revealing that eight tested models suffer performance degradation as noise levels increase despite no information loss.

Original authors: Zain Ul Abedin, Shahzeb Qamar, Lucie Flek, Akbar Karimi

Published 2026-03-17
📖 4 min read☕ Coffee break read

Original authors: Zain Ul Abedin, Shahzeb Qamar, Lucie Flek, Akbar Karimi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a brilliant, super-smart robot assistant that can solve complex math problems. You ask it, "If I bake 8 brownies but need 17, how many more do I need?" The robot thinks for a moment and says, "9!" It's perfect.

Now, imagine you take that same question and sprinkle it with a little bit of "digital glitter"—random punctuation marks like exclamation points, question marks, and semicolons scattered all over the place.

"?! Tiffany baked 8 brownies, but needed 17 ! total for ; her party. If she : used 8 cups of flour..."

You might think, "Well, the words are the same, so the robot should still get it right." But here's the twist: The robot suddenly gets confused and gives the wrong answer.

This is exactly what the paper "ArithmAttack" is about. The researchers wanted to see how fragile these AI math geniuses really are when their instructions get a little messy.

The Experiment: The "Punctuation Prank"

The researchers created a new way to test AI called ArithmAttack. Think of it like a prank where you don't change the story of a math problem, you just ruin the punctuation.

  • The Goal: To see if adding random noise (like !, ?, ;, :) makes the AI fail.
  • The Rule: They didn't delete any words or change the meaning. They just inserted extra marks. It's like taking a perfectly clear sentence and putting a sticker on every third word. The sentence still makes sense to a human, but the AI stumbles.

They tested this on eight different famous AI models (like Llama, Mistral, and DeepSeek) using two popular math test sets (GSM8K and MultiArith).

What They Found: The "Glass House" Effect

The results were surprising. Even though the math problems were the same, the AI models were very sensitive to the noise.

  1. More Noise = More Confusion: Just like how a person might get annoyed and make a mistake if you keep tapping them on the shoulder while they try to think, the AI performed worse as the amount of punctuation noise increased.
  2. The "Smartest" Models Still Stumble: Even the most advanced models (like Llama 3.1) got confused, though they were the least likely to fail completely.
  3. The "Math-Specialist" Wins: One model, Mathstral, was specifically trained on math. It handled the noise better than its "regular" cousin, Mistral. This is like a professional chef handling a messy kitchen better than a home cook; the extra training helps them stay focused.
  4. The Struggler: The model named Zephyr was the most fragile. It was already not great at math, and the punctuation noise made it crash completely.

Why Does This Matter?

You might ask, "Who puts random punctuation in a math problem?"

In the real world, AI doesn't just get perfect, clean text. It gets:

  • Typos from fast typists.
  • Weird formatting from copy-pasting.
  • Noisy data from the internet.

This paper shows that current AI models are like glass houses: they look strong and can solve hard problems, but a little bit of "noise" (like a few extra punctuation marks) can shatter their ability to reason.

The Takeaway

The researchers concluded that while AI is getting better at math, it is still too fragile. It relies too much on the "cleanliness" of the text. If we want AI to be truly reliable in the real world, we need to teach it to ignore the "digital glitter" and focus on the actual meaning of the words, just like a human does.

In short: The paper proves that if you mess up the punctuation enough, even the smartest AI math whiz can turn into a confused mess.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →