← Latest papers
🤖 AI

The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI

This paper argues that optimizing the orchestration layer (the "harness") is more critical than model selection for controlling enterprise Agentic AI costs, demonstrating through a controlled study that a superior harness can reduce token usage by 38% and costs by 41% across diverse models while maintaining or improving task quality.

Original authors: Muayad Sayed Ali, Aliaksandra Novik, Anji Boddupally, Artem Yavorskyi, Chris Nickerson, Daniel Rica, Emily DuGranrut, Felix Leung, Garrett Prince, Grace Barnett, Heath Robinson, Hosain Al Ahmad, Jesse
Published 2026-07-09
📖 4 min read☕ Coffee break read

Original authors: Muayad Sayed Ali, Aliaksandra Novik, Anji Boddupally, Artem Yavorskyi, Chris Nickerson, Daniel Rica, Emily DuGranrut, Felix Leung, Garrett Prince, Grace Barnett, Heath Robinson, Hosain Al Ahmad, Jesse Resnick, Juan Carlos Farah, Jyothi Swaroop Meruga, Leonid Kuznetsov, Luke Gorham, Marie Schmoll, Michael Paciullo, Saumya Das, Sharath Sheripally, Tommy Griscom, Mykyta Osadchyi, Neha Mantri, Nick Westrum, Olivia Benowitz, Parikshith Kulkarni, Radik Chernyshov, Rakshith Vasudev, Rohith Nadimpally, Vikas Gangadevi, Waseem AlShikh

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are running a busy restaurant. In this restaurant, the chefs are the AI models (like Claude, Gemini, or Qwen), and the orders are the tasks you want them to do.

For a long time, the restaurant industry has been obsessed with hiring "super-chefs" who can think deeper and faster. But there's a problem: even with the best chefs, the waiters (the software that manages the orders) have been incredibly inefficient.

This paper, written by a team at Writer, Inc., argues that the real secret to saving money and time isn't just buying a better chef. It's about fixing the waiters. They call this fix the "Harness."

Here is the story of what they found, explained simply:

1. The Problem: "Token Maxing" (The Gluttonous Waiter)

In the world of AI, "tokens" are like the currency you pay for every word the AI reads or writes.

  • The Old Way: Imagine a waiter who, every time a chef needs to think, brings the entire history of the restaurant's day (every order, every conversation, every menu) back to the kitchen. If the chef asks for a second opinion, the waiter brings the whole history again.
  • The Result: The kitchen gets overwhelmed with paper. The chef spends all their time reading old notes instead of cooking. You end up paying a fortune in "tokens" just to keep the waiter busy. The authors call this "Token Maxing": you keep buying more tokens hoping for better results, but you're just paying for the waiter to carry too much weight.

2. The Solution: The "Harness" (The Smart Waiter)

The authors built a new kind of waiter system called the Writer Agent Harness. Instead of dumping everything on the chef, this system is smart about what it brings to the kitchen.

Think of it like a smart filing system:

  • The "Stable" Folder: The waiter keeps the things that never change (like the menu and the chef's identity) in a special, locked folder that the kitchen can glance at instantly without paying extra.
  • The "Volatile" Folder: The things that change every second (like the current time or a specific customer's request) are kept in a small, fresh notepad.
  • The "Summary" Trick: If the conversation gets too long, the waiter doesn't bring the whole book. They write a one-page summary of what happened earlier and only bring that.
  • The "Delegation" Trick: If a task is too big, the main waiter doesn't try to do it all. They send a smaller, specialized helper to do a part of it and just bring back the final answer.

3. The Big Experiment

To prove this works, the authors ran a fair test. They took six different top-tier AI chefs (from different companies) and gave them the same 22 difficult tasks.

  • Group A: Used the "Old Way" (the inefficient waiter).
  • Group B: Used the "Harness" (the smart waiter).
  • The Rule: The chefs were exactly the same. The tasks were exactly the same. The only thing that changed was the waiter.

4. The Results: A Miracle in Efficiency

The results were shocking. By just changing the waiter (the Harness), they got massive improvements across the board:

  • Cost: The bill dropped by 41%. It cost almost half as much to do the same job.
  • Speed: The tasks finished 44% faster.
  • Token Usage: They used 38% fewer tokens.
  • Quality: The quality of the food (the answer) stayed the same, or even got slightly better.

The most surprising part? It didn't matter which chef you hired. Whether you used the most expensive, powerful chef or a cheaper, faster one, the "Smart Waiter" made them all cheaper and faster. In fact, fixing the waiter saved more money than switching from the most expensive chef to the cheapest one would have!

5. The "Harness Leverage"

The paper found a cool pattern: The better the chef was to begin with, the more they benefited from the smart waiter.

  • A super-smart chef + a smart waiter = Amazing results.
  • A weaker chef + a smart waiter = Still good, but they can get a little confused if the waiter tries to do too much complex delegation.

The Bottom Line

For a long time, companies thought the only way to get better AI was to buy more powerful models. This paper says: Stop.

The biggest lever you can pull is how you organize the work. If you build a smart "Harness" that manages the context, caches the boring stuff, and delegates the heavy lifting, you can get the same (or better) results for half the price, regardless of which AI model you use.

It's not about having the biggest engine; it's about having the best transmission.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →