Strategy Executability in Mathematical Reasoning: Leveraging Human-Model Differences for Effective Guidance
This paper identifies a critical gap between the usage of reasoning strategies in successful solutions and their executability when applied as guidance, leading to the proposal of Selective Strategy Retrieval (SSR), a framework that leverages human-model differences to selectively combine strategies and significantly improve mathematical reasoning accuracy across diverse benchmarks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to solve a very tricky math puzzle. You have a brilliant human expert and a super-fast computer robot. Both of them can solve the puzzle correctly, but they do it in completely different ways.
The human expert might say, "Look at the shape of this problem; it's like a triangle, so we should use a special triangle rule."
The robot might say, "Let's just plug all the numbers into a giant formula and crunch the math until we get an answer."
The Problem:
Usually, when we try to help a computer solve a problem, we give it examples of how a human solved similar problems. We assume, "If a human did it this way, the computer should be able to copy it."
But this paper discovered a big surprise: Just because a strategy works for a human doesn't mean the computer can actually do it.
Think of it like this:
- The Human Strategy is like giving a chef a recipe that says, "Add a pinch of 'soul' and 'intuition'." A human chef knows exactly what that means. But if you give that same instruction to a robot chef, it gets confused. It doesn't know how to measure "soul." The strategy exists, but the robot can't execute it.
- The Robot Strategy is like a recipe that says, "Add exactly 3.14 grams of salt." The robot can do this perfectly. But if you give this to a human chef, they might think, "Why so specific? I need to taste it first."
The authors call this gap "Strategy Executability." It's the difference between a strategy being present in a solution and a strategy being doable by the specific model you are using.
The Solution: The "Smart Guide" (Selective Strategy Retrieval)
The researchers built a new system called SSR (Selective Strategy Retrieval). Instead of blindly copying human examples or robot examples, SSR acts like a smart matchmaker.
Here is how SSR works, using a metaphor:
Imagine you are a Travel Agent (the AI model) trying to get a traveler (the math problem) to a destination (the correct answer). You have two guidebooks:
- The Human Guidebook: Full of beautiful, high-level maps and scenic routes.
- The Robot Guidebook: Full of turn-by-turn GPS directions and traffic data.
In the past, travel agents would just pick one guidebook and hope for the best. Sometimes the scenic route was too vague for the GPS, and sometimes the GPS was too boring for the traveler.
SSR is the Super Travel Agent.
Before giving directions, SSR checks:
- "Is this a geometry problem? The Human Guidebook is great for spotting shapes, but the Robot Guidebook is better for calculating angles."
- "Is this an algebra problem? The Robot Guidebook's step-by-step math is perfect here."
SSR doesn't just pick one book. It mixes and matches. It might say, "Use the Human's big-picture idea to start, but switch to the Robot's step-by-step math to finish." It only picks the instructions that the Travel Agent (the AI) is actually capable of following right now.
Why This Matters
The paper tested this on hard math competitions (like the AIME).
- Old Way: Giving the AI a human example improved the score a little bit, but sometimes it made things worse because the AI got confused trying to copy the human's "intuition."
- SSR Way: By picking the executable strategies (the ones the AI can actually do), the AI's accuracy jumped significantly. On some hard tests, it improved by 13 points, which is a huge deal in math competitions.
The Big Takeaway
This paper teaches us that copying isn't always learning. Just because a human does something a certain way doesn't mean an AI should try to copy it exactly.
To make AI smarter, we need to understand how the AI thinks, not just how humans think. We need to give the AI tools it can actually use, rather than tools that look good on paper but are too complex for its "brain" to handle.
In a nutshell:
- The Gap: Humans and AI solve problems differently.
- The Mistake: Assuming AI can copy human methods perfectly.
- The Fix: A system (SSR) that picks the right mix of human and robot strategies based on what the AI can actually do.
- The Result: Smarter, more reliable AI math solving.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.