Unmasking Reasoning Processes: A Process-aware Benchmark for Evaluating Structural Mathematical Reasoning in LLMs
To address the saturation of current benchmarks caused by template-based computation, this paper introduces ReasoningMath-Plus, a process-aware benchmark with 150 structurally complex problems and a corresponding evaluation framework (HCRS and a Process Reward Model) that reveals significant gaps between high answer accuracy and genuine structural reasoning competence in leading LLMs.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a teacher grading a student's math homework.
The Old Way (The "Answer-Only" Test):
In the past, if a student wrote down the correct final number at the bottom of the page, you gave them an "A." You didn't care how they got there. Maybe they guessed, maybe they copied the answer from a cheat sheet, or maybe they made a huge mistake in the first step but got lucky and fixed it later. As long as the final answer was right, the grade was perfect.
The Problem:
Recently, AI models (LLMs) have become so good at these "Answer-Only" tests that they are getting near-perfect scores. But researchers are worried: Are they actually thinking, or are they just memorizing patterns and guessing? It's like a student who memorized the answer key for a test but doesn't understand the math.
The New Solution: REASONINGMATH-PLUS
The authors of this paper built a new, tougher test called REASONINGMATH-PLUS. Instead of just checking the final answer, they want to inspect the student's "thought process" step-by-step.
Here is how they did it, using some creative analogies:
1. The "Skeleton" vs. The "Flesh"
Imagine a human body. The Skeleton is the essential structure that holds everything up. The Flesh is the muscle and skin that makes it look good.
- Existing tests only check if the body looks like a human (the final answer).
- This new benchmark provides a Minimal Reasoning Skeleton for every problem. This is a short, human-written list of the essential logical steps required to solve the problem (e.g., "Step 1: Identify the constraints," "Step 2: Eliminate impossible options").
- The AI is allowed to write as much "flesh" (extra words, explanations, self-corrections) as it wants, but the grader checks if the AI's "skeleton" matches the human-designed one. If the AI misses a crucial bone, the structure collapses, even if the final answer is right.
2. The "Hazard" Score (The Domino Effect)
The paper introduces a scoring system called HCRS (Hazard-aware Chain-based Rule Score).
- The Analogy: Imagine a line of dominoes. If you knock over the first domino incorrectly, the whole chain falls apart. If you knock over the last domino incorrectly, the first 99 were still correct.
- The Innovation: Most grading systems treat every mistake equally. This system realizes that early mistakes are catastrophic. It applies a heavy "Hazard Penalty" if the AI makes a mistake in Step 1 or 2. It's like saying, "You can't fix a broken foundation by painting the roof."
- The Result: Many AI models that got the right answer still got low scores because their reasoning started with a logical error. This exposes "lucky guesses."
3. The Two-Pronged Grader
The researchers used two different "teachers" to grade the AI:
- Teacher A (The Skeletal Auditor): This teacher has the "Gold Skeleton" (the perfect step-by-step plan). They compare the AI's steps directly against this plan. This is strict and precise.
- Teacher B (The Outcome Verifier): This teacher doesn't have the plan. They only see the question and the final answer. They have to guess if the steps make sense just by looking at the result. This simulates a real-world scenario where we don't always have the "answer key" for the thinking process.
- They also trained a smaller AI (a Process Reward Model) to act like Teacher B, learning from Teacher A's corrections so it can grade reasoning without needing the skeleton every time.
The Big Reveal
When they ran the tests, the results were shocking:
- Answer Accuracy: The top AI models got about 5.8 out of 10 on the final answer.
- Reasoning Quality: When graded on their process (using the new Hazard Score), the same models dropped to an average of 4.36 out of 10.
The Conclusion:
The paper argues that we have been fooled by "Answer-Only" metrics. AI models are getting the right answers, but often through fragile, illogical, or memorized paths. By looking under the hood at the structural reasoning, we see that their true "thinking" ability is much weaker than we thought.
In short: This paper is like a mechanic who stops just checking if a car's engine starts (the answer) and instead inspects the wiring, the fuel lines, and the timing (the reasoning process) to see if the car will actually run safely on a long trip.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.