Compression, structure, and executor capability: a controlled real-cost decomposition of language-model agent skill optimisation
This controlled study of 1,200 language model agent rollouts demonstrates that while various skill optimization strategies fail to improve cost-efficiency or performance, upgrading the executor model remains the only factor that significantly boosts task success rates, albeit at a substantially higher monetary cost.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a team of digital assistants (AI agents) to fix complex software problems. You have a "skill manual" for them—a set of instructions on how to solve specific tasks. The big question researchers asked is: How should we write and deliver these manuals to get the best results for the lowest price?
Many people assume that if you make the manual shorter, organize it into fancy charts, or use a super-smart editor to rewrite it, the assistant will work faster and cheaper. This study puts that assumption to the test.
Here is what they found, explained simply:
1. The Experiment: The "Cooking" Analogy
Think of the AI agent as a chef.
- The Skill: The recipe the chef follows.
- The Compiler: The person who edits the recipe (maybe shortening it or turning it into a flowchart).
- The Executor: The actual chef cooking the meal.
- The Cost: The price of the ingredients and the chef's time.
The researchers tested 10 different ways to serve the recipe to the chef. They tried:
- Giving the chef the original, unedited recipe.
- Cutting the recipe down to just the essentials (shortening).
- Rewriting the recipe into a structured list or a fancy table (structured rendering).
- Only showing the chef the part of the recipe relevant to the current dish (scoped loading).
- Using a "Junior Chef" (a cheaper, smaller AI model) vs. a "Master Chef" (a more expensive, powerful AI model) to cook.
They ran this experiment 1,200 times across 40 different software tasks, measuring two things: Did the dish come out right? (Success rate) and How much did it cost in real money?
2. The Big Surprise: The "Chef" Matters More Than the "Recipe"
The most important finding is that who does the cooking matters way more than how the recipe is written.
- The Master Chef Wins: When they used the powerful "Master Chef" (the stronger AI model), the success rate jumped by 27%. However, this cost about 5 times more per task.
- The Junior Chef Struggles: When they used the cheaper "Junior Chef," the fancy recipe tricks didn't help. In fact, they often made things worse.
- Shortening the recipe: Didn't really help the Junior Chef cook better, and it didn't save money.
- Fancy formatting (charts/tables): Actually confused the Junior Chef, lowering their success rate compared to a simple, plain-text recipe.
- Showing only parts of the recipe: Also lowered success rates without saving money.
The Metaphor: Imagine you have a student (the Junior Chef). If you give them a simplified, bullet-pointed study guide, they might still fail the test if they don't understand the basics. But if you hire a genius tutor (the Master Chef), they will pass the test easily, even if the study guide is messy and long. The quality of the worker is the deciding factor, not the format of the instructions.
3. The Cost Trap: "Token Count" vs. "Real Money"
There was a trap in how people usually measure AI costs.
- The Token Illusion: People often count "tokens" (chunks of text) to estimate cost. They thought, "Oh, the fancy structured recipe uses fewer tokens, so it's cheaper!"
- The Reality: The "Master Chef" uses fewer tokens because they are smarter and need less instruction, but they charge a much higher price per token.
- Analogy: It's like hiring a master carpenter who finishes a job in 1 hour but charges \500/hour, versus a novice who takes 5 hours at \20/hour. Even though the master uses less "time" (tokens), the total bill is much higher.
- The study found that when you look at the actual dollar cost, none of the "optimized" recipes (shortened, structured, or scoped) actually saved money compared to just giving the raw, original instructions.
4. The "Break-Even" Myth
Some people think: "If I pay a little extra now to have a super-smart editor rewrite the recipe, I'll save money later because the chef will work faster."
- The Verdict: The math says no.
- Analogy: Imagine paying \10 to have a professional editor rewrite your instructions. You hope this saves you \1 every time you use them. The study found that for most tasks, you would have to use that rewritten recipe hundreds of times just to break even. Since most people only use these skills a few times, you are almost always losing money by trying to "optimize" the instructions.
Summary of Findings
- Don't over-engineer the instructions: Making the instructions shorter, structured, or "scoped" didn't make the AI smarter or cheaper. In fact, for the cheaper AI models, it often made them worse.
- The raw material is fine: The original, unedited "skill" instructions worked just as well (or better) than the fancy versions.
- Upgrade the worker, not the manual: The only thing that reliably improved results was switching to a more powerful AI model (the "Master Chef"), even though it costs more.
- Real cost matters: Don't be fooled by "token counts." Always look at the real dollar price. In this study, trying to save on tokens by rewriting instructions didn't actually save money.
The Bottom Line: If you want your AI agent to succeed, stop worrying about how pretty or short your instructions are. Instead, focus on giving the task to a more capable AI model. The "quality of the worker" is the only lever that truly moves the needle.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.