← Latest papers
🤖 AI

Scaffold Effects on GAIA: A Controlled Comparison

This pre-registered controlled study demonstrates that scaffold choice significantly impacts agent performance on GAIA benchmarks, revealing that measured capability scores are highly scaffold-conditional and that the anticipated reduction in scaffold sensitivity as models improve does not hold, with structured multi-agent designs often yielding substantial accuracy gains over standard ReAct approaches.

Original authors: Jason Starace

Published 2026-06-09
📖 4 min read☕ Coffee break read

Original authors: Jason Starace

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to judge how good a chef is at cooking a complex meal. You have three different kitchens (scaffolds) to test them in:

  1. The Solo Chef (ReAct): The chef thinks, grabs an ingredient, cooks, thinks again, and grabs the next ingredient, all in one continuous stream.
  2. The Kitchen Team (Planner-Actor-Rater): A manager writes a recipe, a cook executes the steps, and a critic tastes the dish after every step to make sure it's right before moving on.
  3. The Blueprint & Builder (Planner-then-Executor): First, a planner writes a detailed, step-by-step blueprint with no tools. Then, a builder takes that blueprint and builds the dish.

This study asked a simple question: Does the kitchen you put the chef in change how good they seem to be?

The answer is a resounding yes. In fact, the "kitchen" matters just as much as the chef's actual talent.

The Big Discovery: The "Kitchen" Matters More Than You Think

The researchers tested five different "chefs" (AI models from Anthropic, Google, and OpenAI) on a set of difficult puzzles (called GAIA). They found that simply switching the kitchen setup could change a model's score by 28 percentage points.

To put that in perspective: If you tested a model in a "Solo Chef" kitchen and it got a 56% score, but then you put the exact same model in a "Kitchen Team" setup, it might suddenly get an 84% score. The chef didn't change; the way they were allowed to work did.

This means that when we see headlines like "Model X is 80% smart," that number isn't just about the model. It's a mix of the model's brain plus the specific instructions and tools it was given. If you compare two models tested in different kitchens, you aren't comparing their brains; you're comparing their kitchens.

Myth Busting: "Smarter" Models Don't Need Less Help

The researchers had a hunch (hypothesis) that the most powerful, "frontier" models would be so smart that they wouldn't care much about the kitchen setup. They thought a super-brain could handle a messy kitchen just as well as a clean one.

They were wrong.

In fact, the most powerful model they tested (Anthropic's Opus) actually gained the most from having a structured, organized kitchen (the "Kitchen Team" setup) when the tasks were hard. The "smartest" model wasn't the most independent; it was the one that benefited most from having a manager and a critic. The gap between the best and worst performance didn't shrink for the smart models; it stayed wide or even got bigger.

The "Family" vs. "Rank" Surprise

Another surprise was that the benefit of a fancy kitchen depended on which family the model came from, not just how "ranked" it was.

  • The Anthropic family (Haiku, Sonnet, Opus) all got a huge boost from the "Kitchen Team" setup.
  • The models from Google and OpenAI didn't get that same boost.

It's like saying that a specific type of car engine runs much better with a specific brand of fuel, but a different brand of engine doesn't care. You can't just assume a "higher tier" model will always benefit more from better tools; it depends on who made the engine.

Cost and Efficiency

The study also looked at the "receipt" (cost).

  • The "Solo Chef" (ReAct) was the most expensive and slowest. It made a lot of mistakes and had to try again often.
  • The "Kitchen Team" and "Blueprint" setups were faster, made fewer mistakes, and recovered better when things went wrong mid-task.
  • The Winner: A specific combination of Google's Gemini model with the "Blueprint" setup was the cheapest and most accurate option for the hardest tasks.

The Bottom Line

This paper tells us that capability numbers are conditional. You cannot say a model is "X% capable" in a vacuum. You have to say, "X% capable when using this specific set of instructions and tools."

If you want to know how good an AI really is, you can't just look at one score. You have to realize that the score is a snapshot of the model plus the scaffolding holding it up. If you change the scaffolding, the score changes, sometimes dramatically. The "gap" between what a model can do and what it actually does in a test is huge, and it doesn't necessarily get smaller just because the model gets smarter.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →