SKIMIX: Multi-Agent Harness-Time Scaling with Skill Mixture for Dynamic Harness Engineering
The paper introduces SKIMIX, a multi-agent framework that leverages collaborative skill refinement and adaptive routing to significantly enhance open-ended mathematical reasoning, while revealing that its benefits are task-dependent and non-monotonic with respect to agent count.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to solve a really tricky puzzle, like a complex math problem or a mystery. In the world of artificial intelligence, we have built "agents"—smart computer programs that can think and act. To make these agents super helpful, we give them a "harness," which is like a giant backpack filled with different tools and skills. Some tools are for searching the web, some are for writing code, and some are for breaking down big problems into smaller steps. The big question researchers are asking right now is: if we give an agent a backpack with hundreds of different tools, how do we make sure it picks the right ones? If we just throw everything at the problem, the agent might get confused by too many choices, or worse, it might pick the wrong tool and get stuck. This paper explores a clever new way to manage these toolkits so that teams of AI agents can work together without tripping over each other.
The researchers behind this study, Jia Luo from Huazhong University of Science and Technology, propose a new system called SKIMIX. Think of SKIMIX not as a single genius trying to do everything alone, but as a lively team of detectives, each carrying a slightly different version of that giant backpack. Instead of one detective trying to figure out which tools to use, the team splits up. One detective grabs the "math tools," another grabs the "search tools," and a third grabs the "logic tools." They all look at the same mystery, write down their best guesses, and then share their notes with each other. They do this over and over, refining their answers in rounds, like a group of friends whiteboarding a solution together.
The most exciting part of their discovery is that this teamwork doesn't always work the same way. It depends entirely on the type of puzzle they are solving. When the task is an open-ended math problem (like creating a solution from scratch), having a diverse team is a game-changer. On a tough math test called AIME, a single agent trying to fix its own mistakes got the answer right 40% of the time. But when they used the SKIMIX team with just three different agents, the success rate jumped to 73.3%. With fifteen agents, it went even higher to 76.7%. The variety of approaches helped them find the right path much faster.
However, the paper also reveals a surprising twist: more agents don't always mean better results. In fact, for multiple-choice questions (where you just have to pick the right answer from a list of options), the team approach actually made things worse. On a science quiz called GPQA Diamond, a single agent refining its own thoughts got 88% of the answers right. But when they added more agents with different toolkits, the score dropped. With three agents, it fell to 78%; with fifteen, it plummeted to 72%. It seems that for multiple-choice questions, having too many different opinions just creates noise and confusion, whereas a single, focused thinker does the best job.
The researchers also found that the "sweet spot" for teamwork happens very quickly. Most of the improvement happens in the second round of sharing ideas. By the third round, the team often stops getting better and might even start making mistakes, suggesting that you don't need to keep the meeting going forever. They also noticed that while the team often had the right answer hidden somewhere in their notes (high "coverage"), they sometimes failed to pick it as the final answer (lower "accuracy"). This suggests that while the team is great at generating ideas, they still need a better way to vote on the final winner.
In short, SKIMIX is a smart system that manages a mix of AI skills to prevent "skill dilution"—a fancy way of saying it stops the agents from getting overwhelmed by too many useless tools. The study suggests that we shouldn't just throw more agents at every problem. Instead, we should be strategic: use a diverse team for open-ended, creative challenges, but stick to a single, focused agent for multiple-choice tests. It's a reminder that in the world of AI, sometimes less is more, and the right mix of tools matters more than the size of the toolbox.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.