When Do LLM Agents Treat Surface Noise Differently from Semantic Noise? A 68-Cell Measurement Study with a Held-Out Trace-Level Validation
This empirical study of 68 experimental cells across ten large language models demonstrates that meaning-bearing perturbations disrupt agent reasoning and final answers significantly more than presentation-based noise of comparable severity, a finding validated on a held-out model and attributed to a "stealth-divergence" mechanism where semantic noise induces divergence in intermediate reasoning steps rather than initial actions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, over-achieving student (an AI Agent) who solves complex math and logic problems by talking to themselves out loud. This student writes down every step of their thinking before giving the final answer. This is called "Chain of Thought."
The researchers in this paper wanted to know: Does it matter how you ask the student a question?
Specifically, they tested two types of "noise" or changes to the question:
- The "Meaning" Change: Rewording the question using synonyms or paraphrasing it (e.g., changing "How many apples?" to "What is the total count of fruit?"). The meaning is the same, but the words are different.
- The "Presentation" Change: Shuffling the order of words, changing the font, adding random distracting sentences, or messing up the spacing. The meaning is the same, but the look is different.
The Big Discovery: The "Silent Corruption"
The researchers found a surprising gap. When they changed the words (Meaning Change), the student was much more likely to get the wrong answer at the end, even though the question meant the same thing. When they just changed the format (Presentation Change), the student was much more likely to get it right.
It's as if the student is very sensitive to what you say, but not so sensitive to how you say it.
How They Tested It
They didn't just guess; they ran a massive experiment:
- The Class: They used 10 different "students" (AI models) from 7 different families (like Llama, Qwen, Mistral, etc.).
- The Homework: They gave them 3 different types of tests (Math, Logic, and Reading Comprehension).
- The Volume: They took 1,530 original questions and created about 11,000 variations of them.
- The Stress Test: They made sure the "Meaning" changes and "Presentation" changes were equally difficult to handle (so they could compare them fairly).
The Result: On average, the "Meaning" changes caused the AI to fail 20% more often than the "Presentation" changes. This happened in 64 out of 68 different test scenarios.
The "Stealth Divergence" Mystery
Here is the most interesting part. The researchers looked at the student's step-by-step notes (the "trace") to see when the mistake happened.
They expected that if you change the meaning of the question, the student would get confused immediately and start writing nonsense right away.
- What they found instead: The student started correctly! The very first step of their thinking was identical for both types of changes.
- The Twist: Starting from Step 2, the student's internal thoughts began to drift apart. The "Meaning" changes caused a "silent corruption." The student kept writing, but their internal logic slowly became confused, leading to a wrong answer at the end.
- The Metaphor: Imagine two cars driving down the same road.
- Presentation Noise: The road signs are upside down or the paint is peeling. The driver notices immediately, stops, and corrects their path.
- Meaning Noise: The road signs look perfect, but the map inside the driver's head has a tiny, invisible smudge. The driver starts driving fine, but by the second turn, they are subtly heading the wrong way, and by the end, they are lost. They never realized they were off course until it was too late.
What They Didn't Find (The Retractions)
Science is about being honest, and the authors admitted they were wrong about two things they thought they knew:
- They thought the "Meaning" changes would make the student get confused faster (at Step 1). They were wrong; the confusion started at Step 2.
- They thought the student would try to "fix" their own mistakes less often with "Meaning" changes. They were wrong; the fixing rate was the same.
They also found that this "Meaning vs. Presentation" gap isn't the same for every AI. If you swap the tool used to generate the questions, the ranking of which AI is "most sensitive" changes. So, you can't say "AI Model X is always 20% worse than Model Y" without knowing exactly how the test was set up.
The Bottom Line
This paper is a measurement study. It proves that AI agents are surprisingly fragile when you rephrase a question, even if the meaning stays the same. They don't break immediately; they break slowly and silently, starting from the second step of their thinking process.
The authors released all their data and tools so other researchers can check their work, but they warn that this is a diagnostic tool for researchers, not a magic switch for fixing AI in the real world yet.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.