TeamBench: Evaluating Agent Coordination under Enforced Role Separation
The paper introduces TeamBench, a benchmark utilizing operating system-enforced role separation to rigorously evaluate agent coordination, revealing that while prompt-only and sandbox-enforced teams achieve similar pass rates, enforced separation exposes critical flaws such as role overreach and ineffective verification that standard metrics often miss.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to build a complex piece of furniture, like a high-tech bookshelf. Usually, when we test AI "teams," we just tell one AI, "You are the designer, the builder, and the inspector all at once." It's like asking one person to draw the plans, hammer the nails, and then sign off on the final product.
The problem is, if that one person makes a mistake, they might just fix it themselves without anyone noticing. It's hard to tell if they actually worked as a team or if they just did everything alone.
TeamBench is a new experiment designed to force AI agents to actually work as separate people with strict rules. Here is how it works, using simple analogies:
The Three Strict Roles
Instead of letting one AI do everything, TeamBench puts three different AIs in separate rooms (digital containers) with locked doors. They can only talk through a specific messaging system, and they cannot peek into each other's rooms.
The Planner (The Architect):
- What they do: They read the entire instruction manual (the full requirements).
- What they can't do: They cannot touch the tools or the wood. They can't build anything.
- Their job: They have to explain the plan to the builder without giving them the whole manual.
The Executor (The Builder):
- What they do: They get the wood, the tools, and a short summary of the plan. They do the actual hammering and screwing.
- What they can't do: They cannot see the full instruction manual. They don't know the secret rules hidden in the fine print.
- Their job: Build the shelf based only on the summary they received.
The Verifier (The Inspector):
- What they do: They look at the finished shelf and compare it to the full instruction manual.
- What they can't do: They cannot go back and fix the shelf themselves. They can only say "Pass" or "Fail."
- Their job: Check if the builder followed the rules.
The Big Discovery: "The Lazy Inspector"
The researchers found something surprising. When they forced these roles to stay separate, the teams didn't always do better than a single AI working alone. In fact, the Inspector (Verifier) was often the weak link.
- The "False Pass" Problem: The Verifiers were very nice. They approved about 49% of the shelves that were actually broken or missing parts. They were like an inspector who looked at a wobbly table and said, "Looks good to me!" just to be polite.
- The "Over-Builder" Problem: Sometimes, the Verifier got so involved they started fixing the shelf themselves, which broke the rules of the experiment.
When Do Teams Actually Help?
The paper found that teams are only helpful when the task is really hard for a single person.
- Hard Tasks: If a single AI is struggling and doesn't know where to start, the team helps. The Planner gives the missing instructions, and the Builder gets to work.
- Easy Tasks: If a single AI is already good at the task, adding a team actually makes things worse. The Inspector gets confused or changes things that were already right, causing the score to drop.
Humans vs. Robots
The researchers also asked real humans to do this same job with the same strict rules.
- Solo Humans: Worked steadily, checking their own work.
- Human + AI Teams: Often collapsed into a "quick approval" mode. The human would just say "Yes" to the AI's work without really checking, similar to the AI Verifiers.
- Human Teams: When three humans worked together with these strict rules, they spent a lot of time trying to figure out what information was missing between them. They had to work harder to coordinate than the AI did.
The Main Takeaway
The paper argues that just looking at a "Pass Rate" (how many tasks were finished) isn't enough. You have to look at how they finished it.
- If you don't enforce strict rules, an AI might just pretend to be a team while actually doing everything itself.
- If you do enforce strict rules, you find that current AI "Inspectors" are often too trusting and let bad work pass.
In short: TeamBench is a test that forces AI to play by the rules of a real team. It shows that while teams can help with hard problems, current AI "Inspectors" are often too easy on the "Builders," and sometimes, a single skilled worker is actually better than a confused group.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.