← Latest papers
💬 NLP

Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation

This study demonstrates that controlled lexical and syntactic perturbations significantly degrade the performance and destabilize the rankings of large language models across major benchmarks, revealing a reliance on surface-level patterns over robust linguistic competence and highlighting the critical need for robustness testing in standard evaluations.

Original authors: Bogdan Kostić, Conor Fallon, Julian Risch, Alexander Löser

Published 2026-02-20
📖 5 min read🧠 Deep dive

Original authors: Bogdan Kostić, Conor Fallon, Julian Risch, Alexander Löser

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a judge in a cooking competition. You have 23 different chefs (the AI models) and you want to see who makes the best dish. Usually, you give them the exact same recipe and ingredients, taste the result, and rank them from 1st to 23rd place.

This paper is about what happens when you secretly tweak the recipe just a tiny bit—without changing the actual taste or the final dish—and see if the judges (the AI) get confused.

Here is the story of their experiment, broken down simply:

1. The Setup: The "Same Dish, Different Menu"

The researchers took three famous "recipe books" (benchmarks) used to test AI:

  • MMLU: A general knowledge quiz (like a trivia night).
  • SQuAD: A reading comprehension test (finding answers in a story).
  • AMEGA: A medical guideline test (checking if a doctor follows the rules).

They asked 23 different AI chefs to solve these tests. But then, they created two special "tricks" to see if the AI was actually smart or just memorizing the menu:

  • Trick #1: The Synonym Swap (Lexical Perturbation).
    Imagine the recipe says, "Add a large pinch of salt." The researchers changed it to, "Add a big pinch of salt." The meaning is identical, but the words are different.
  • Trick #2: The Sentence Shuffle (Syntactic Perturbation).
    Imagine the recipe says, "The chef chopped the onions." They changed it to, "The onions were chopped by the chef." The action is the same, but the sentence structure is flipped.

2. The Big Surprise: The AI is a "Word-Matcher," Not a "Thinker"

When the researchers ran the tests, they found something shocking:

  • The AI hates new words: When they swapped words (like "large" to "big"), the AI's performance crashed. It was like a student who memorized the answer key but failed the test because the teacher wrote "big" instead of "large." The AI seemed to rely heavily on specific words it had seen before, rather than understanding the concept.
  • The AI is okay with sentence flips: When they rearranged the sentence structure, the AI was much more stable. It could handle the shuffle better than the word swap.

The Analogy: Think of the AI like a parrot. If you teach a parrot to say, "The sky is blue," it might get confused if you ask, "Is the sky blue?" or "Blue is the sky." It's reacting to the specific sounds it heard, not the idea of the sky. This paper shows that current AIs are acting more like parrots than deep thinkers.

3. The Leaderboard is a "House of Cards"

In the AI world, there are public leaderboards (like a high-score list) that tell us which AI is the "best."

The researchers found that these leaderboards are brittle.

  • On the general knowledge test (MMLU), the rankings stayed mostly the same.
  • But on the reading and medical tests, the rankings flipped wildly.
    • A model that was ranked #1 might drop to #12 just because someone changed a few words.
    • A model that was ranked #14 might jump to #9 just because the sentence structure changed.

The Metaphor: Imagine a race where the runners are ranked by who crosses the finish line first. But then, the referee changes the color of the finish line tape from red to blue. Suddenly, the runner who was second wins, and the winner loses. The paper says our current AI leaderboards are like that race—they are unstable and can't be trusted if the "tape color" (the wording) changes slightly.

4. Bigger Isn't Always Better

There is a common belief in tech that "bigger models are smarter and tougher." The researchers tested this by comparing tiny AIs to massive ones.

  • The Result: Size doesn't guarantee safety.
    • On the trivia test, the biggest models actually got the worst scores when the words were changed. They were so specialized that they broke easily.
    • On the reading test, the bigger models were actually more robust.

The Analogy: It's like a bodybuilder. A bodybuilder might be incredibly strong at lifting heavy weights (big model on a specific task), but if you ask them to do a delicate dance move (a different task), they might be clumsier than a lightweight gymnast. Being "big" doesn't make you good at everything.

The Takeaway: What Should We Do?

The paper concludes that we are currently overestimating how "smart" these AIs are. They are very good at spotting patterns in the words they've seen before, but they aren't truly understanding the meaning.

The Solution:
Before we trust an AI with important jobs (like diagnosing diseases or writing legal contracts), we shouldn't just ask, "What is your score?" We should ask, "How do you handle it if I say the same thing in a different way?"

We need to start testing AI for robustness (toughness) just as much as we test them for accuracy (correctness). If an AI breaks when you change a synonym, it's not ready for the real world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →