FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks
This paper introduces FormInv, a measurement protocol that exposes how semantic invariance flaws in mathematical reasoning benchmarks distort model rankings and reveals a critical gap between aggregate accuracy and semantic consistency, demonstrating that benchmark design choices implicitly determine which models appear superior.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: The "Shape-Shifting" Test
Imagine you have a student taking a math test. You ask them, "Is the square root of 4 greater than or equal to 0?" They answer, "Yes." Then, you ask the exact same question but change the wording slightly: "Is 0 less than or equal to the square root of 4?"
If the student is truly smart, they should answer "Yes" to both. But what if they get the first one right and the second one wrong?
That is exactly what this paper discovered about the most advanced AI models (like GPT-4o, Claude, and DeepSeek). Even though they are incredibly good at math, they are surprisingly fragile. If you change the shape of the question without changing the meaning, the AI often gets confused and gives a different answer.
The authors call this lack of stability a "Semantic Invariance Gap." In plain English: The AI doesn't understand the math; it just recognizes the pattern of the sentence.
The Problem: The "Trick Question" Trap
The paper argues that current benchmarks (standard tests for AI) are flawed because they only ask questions in one specific way. It's like testing a driver only on a straight highway. They might drive perfectly there, but if you ask them to drive the same route but with the lanes swapped, they might crash.
The researchers found that:
- High scores are misleading: An AI might get 96% of the answers right on a standard test. But when you ask the same questions in different ways, that "96%" drops significantly because the AI fails to realize the questions are the same.
- The "Ranking Reversal": This is the most shocking part. Depending on how you ask the questions, the winner changes.
- If you ask questions in "Order A," Model X wins.
- If you ask the same questions in "Order B," Model Y wins.
- It's like a sports league where the team that wins the championship changes every time you switch the rules of the game, even though the players are the same.
The Solution: The "Unanimity Audit"
How do you catch these bad questions? The authors created a protocol called FormInv.
Think of it like a panel of judges at a talent show. If you have 9 different judges (AI models) and they all agree that a specific question is confusing or broken, it's likely the question is broken, not the judges.
- The Protocol: They take a question and ask 9 different AI models to answer it.
- The Discovery: If 6 out of 9 models get a "paraphrased" (reworded) version of the question wrong, but they all get the original version right, the system flags the reworded question as broken.
- The Result: They found that about 3% to 47% of the "reworded" questions in existing tests were actually mathematically incorrect or confusing. The AI wasn't failing the math; the test was failing the AI.
The "No-Free-Benchmark" Rule
The paper makes a bold claim: There is no perfect test.
Imagine you are trying to rank 9 different cars.
- If you test them on speed, Car A wins.
- If you test them on fuel efficiency, Car B wins.
- If you test them on handling, Car C wins.
The authors say the same thing happens with AI math tests. Because no single AI is perfect at every type of rewording, the person designing the test gets to choose which "type" of question to include. By choosing the questions, they are secretly deciding who wins the race. The paper provides a tool to make this choice transparent so we know why a model won.
The Key Takeaways
- Accuracy isn't everything: A model can be 96% accurate but still fail half the time when you rephrase a question. This means it doesn't truly "understand" the math; it's just guessing based on patterns.
- The "Invariance Gap": This is a new score that measures how consistent an AI is. If an AI gets the same answer for the same question regardless of how it's asked, it has a high "Invariance" score. The best models in the study still only got about 82% consistency, meaning they were inconsistent on nearly 1 in 5 questions.
- Fixing the Tests: The authors built a tool (FormInv) that automatically checks if a test question is "broken" by seeing if multiple AI models get confused by it. They used this to fix existing tests, which changed the ranking of the top AI models.
Summary Analogy
Imagine you are testing a translator.
- Old Way: You ask the translator to translate "The cat is on the mat" into French. They do it perfectly. You give them an A.
- New Way (FormInv): You ask them to translate "The mat has a cat on it," "Is the cat on the mat?" and "On the mat sits the cat."
- The Result: The translator gets the first one right but fails the others. You realize they don't actually understand French; they just memorized the first sentence.
FormInv is the tool that forces us to ask the second set of questions so we can see who truly understands the language and who is just memorizing patterns.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.