Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions
This paper introduces an exam-style evaluation framework revealing that current reasoning models fail to strategically allocate a shared test-time compute budget across multiple questions, instead behaving as greedy sequential solvers that prioritize presentation order over question difficulty or value.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are sitting in a giant, high-tech library where the books are written by super-smart computers called "reasoning models." These computers are amazing at solving puzzles, but they have a strange quirk: the harder the puzzle, the more they "think" before answering. This thinking process costs something called "compute," which is like a limited supply of energy or a strict time limit on a test. Scientists have long known that if you give these computers more time to think on a single hard problem, they get better at it. But a new question has popped up: What happens when you give the computer a whole stack of different puzzles, but only one shared time limit for the entire stack? Can the computer decide which puzzles are worth spending its energy on, or does it just plow through them in the order they appear, wasting its time on the easy ones and running out of steam before it even sees the hard ones? This is the big mystery researchers are trying to solve, because if these smart computers can't manage their own energy across a group of tasks, they might fail in real-world situations where they need to juggle many jobs at once.
This paper, titled "Thinking Hard, Not Smart," dives into that exact problem. The researchers set up a digital "exam" for several of the world's most advanced reasoning models. Instead of giving each question its own time limit, they gave the whole exam a single, shared budget of "thinking tokens" (a unit of computer effort). The exam included questions of varying difficulty, some worth more points than others, and the models had to figure out how to split their limited energy to get the highest total score. The results were a bit of a shock: the models failed to act like strategic exam-takers. Instead of skipping a hard, low-point question to save energy for an easy, high-point one, they behaved like a greedy student who just reads the questions from top to bottom. They spent a huge amount of effort on the first few questions, often running out of tokens before they even got to the end of the paper, regardless of how many points the later questions were worth.
The study found that these models are "position-blind" strategists. They don't seem to care if a question is labeled "Hard" or "Easy," or if it's worth 20 points or 5. They mostly just follow the order the questions are presented. If a question is first, they pour all their energy into it; if it's last, they often ignore it completely. The researchers tested this by shuffling the order of the questions and changing their point values, but the models didn't adapt. Even when the researchers gave the models a special prompt telling them to "plan ahead" and "allocate wisely," the models didn't change their habits. They still spent their energy in the same way, just spreading the same lack of strategy a bit more evenly. The paper suggests that while these models are getting better at "thinking hard" on a single problem, they haven't yet learned how to "think smart" about which problems are worth thinking about.
To make this clearer, imagine you have a backpack with a strict weight limit of 10 pounds, and you are at a flea market with 20 items. Some items are heavy but worth a lot of money (high value, high cost), some are light but worth very little (low value, low cost), and some are heavy but worthless. A smart shopper would look at the whole list, pick the items that give the best "value per pound," and fill the bag to get the most money. But these AI models acted like a shopper who just grabs the first item they see, then the second, then the third, stuffing them in until the backpack is too heavy to lift. By the time they reached the back of the line, they had no room left for the best deals. The researchers found that as the number of questions (or items) grew from 5 to 20, the models got worse at covering the whole list. For example, with 20 questions, the models only did serious work on about 4 to 8 of them, leaving the rest completely untouched.
The paper also looked at whether the models were reacting to the difficulty of the questions. It turns out they do spend more time on hard questions, but only after they have already started them. They don't look at a hard question and say, "This looks tough, I'll skip it to save energy for later." Instead, they dive in, realize it's hard, and keep going until they run out of energy. This is what the authors call "reactive" rather than "prospective" thinking. They are good at solving a problem once they start, but bad at deciding which problems to start in the first place.
Interestingly, the researchers tested this on both math problems and coding problems, and the results were the same. Whether the task was solving an equation or writing a computer program, the models behaved like a train on a fixed track: they moved forward, one question at a time, until the fuel ran out. They didn't jump tracks to the easier or more valuable stops. The study concludes that this ability to manage a shared budget across multiple tasks is a distinct skill that current models lack. They know how to think hard, but they haven't figured out how to ration their thinking. This isn't just a small glitch; it suggests that as we ask these models to do more complex, multi-step jobs, we might need to teach them a new kind of "metacognition"—the ability to think about their own thinking and manage their resources, not just solve the puzzle in front of them.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.