Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility
This paper introduces **CostAda**, a cost-calibrated adaptive controller that optimizes LLM discovery by dynamically allocating search resources based on the ratio of quality gain to token cost, thereby achieving superior or equivalent results with significantly reduced budgets compared to existing methods.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a team of brilliant detectives trying to solve a series of incredibly complex puzzles. They have a limited supply of "thinking tokens"—a magical currency that pays for every word they write, every clue they search for, and every hypothesis they test. In the world of Artificial Intelligence, these tokens are the real cost of using Large Language Models (LLMs) to discover new scientific formulas, design better algorithms, or write perfect computer code. For a long time, the "controllers" (the managers of these detective teams) only cared about one thing: Did the score go up? If a detective solved a tiny clue or a massive mystery, and both gave the same score boost, the manager treated them exactly the same. But here's the catch: solving the tiny clue cost one token, while solving the massive mystery cost six. If the manager kept paying for expensive moves just because the score ticked up, the team would run out of money long before they found the best solution. The big question for scientists is: How do you manage a team when every move costs a different amount, and you have a strict budget to stay within?
This paper introduces a new, smarter manager called CostAda. The researchers found that the old way of managing—ignoring how much a move actually cost—was a recipe for disaster. They proved mathematically that if you don't adjust your credit based on the price tag, you might end up with almost zero progress as the search gets more complicated. CostAda changes the game by looking at the "bang for the buck." It asks: "Is this score improvement worth the tokens we just spent?" If a move is expensive, CostAda demands a bigger reward before it approves the next step. It also keeps a close eye on the remaining budget. Early in the game, when money is plentiful, it's okay to take a few expensive risks to find a new path. But as the budget runs low, the manager becomes strict, only approving moves that are cheap and proven to work, or saving the last few tokens for a big, game-changing guide.
The team tested this new manager on twelve different puzzle sets (benchmarks) using two different AI brains (GLM-5 and GPT-5.4). The results were impressive. In twelve out of sixteen test cases, CostAda managed to find solutions just as good as the best previous methods, but it did so using at most half the budget. In other words, it got the same high-quality results for 50% of the price. Across all eight major challenges, CostAda consistently achieved the highest average final quality. The researchers showed that by treating the cost of every action as a critical part of the decision-making process—rather than just a receipt to look at later—the AI can explore more efficiently, avoid wasting money on dead ends, and reach the finish line faster and smarter. It's like realizing that sometimes, taking the expensive highway is a waste of gas, but knowing exactly when to pay for it to get there first is the key to winning the race.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.