MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers
This paper introduces MCPEvol-Bench, a novel benchmark that evaluates LLM agent adaptability to dynamic toolset evolutions in MCP servers by simulating realistic interface changes, revealing that even state-of-the-art models struggle with such adaptations and suffer significant performance declines.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are building a super-smart robot assistant. You teach it to use tools—like a calculator, a calendar, or a search engine—to solve problems for you. In the world of artificial intelligence, these tools are often connected through a standard system called the Model Context Protocol (MCP), which acts like a universal plug-and-play adapter, letting your robot grab any tool it needs from a vast digital toolbox. For a long time, scientists tested these robots by giving them a static set of tools and seeing if they could solve a puzzle. It was like testing a driver on a track where the road signs never moved and the traffic lights never changed. But in the real world, tools are messy and alive; they get updated, renamed, or replaced constantly. The big question is: if the toolbox changes while the robot is working, will it get confused and crash, or will it adapt and keep going?
This is exactly what the researchers behind MCPEvol-Bench wanted to find out. They realized that previous tests were too easy because they didn't account for the fact that tools evolve. To fix this, they built a new, dynamic testing ground called MCPEvol-Bench. Instead of a static track, they created a "living" environment where the tools the robots use change right in the middle of the task. They studied 123 real-world tool servers and found that in the real world, about 20% of remote servers go offline, and more than half of the tools inside them get deleted or replaced over time. To mimic this chaos, they invented 11 "mutation operators"—think of them as digital mutation viruses—that automatically tweak the tools. They might add a new button, rename a parameter, or delete a feature entirely, just like a real software developer would do during an update.
The results were a bit of a wake-up call. The researchers tested 12 of the most advanced AI models available, including heavy hitters like GPT-5.4 and Claude-Sonnet-4-6. In the beginning, when the tools were stable, these models performed well. But as soon as the tools started evolving, the robots stumbled. Even the smartest models saw their success rates drop significantly—by about 13.7% to 14.4%—when the tools changed. The study found that the robots didn't necessarily forget how to use the tools; instead, they got lost in the planning and reasoning stages. It's like a chef who knows how to chop onions perfectly but suddenly finds the knife has been replaced with a spoon; they don't know how to adjust their recipe, so they end up making a mess. The researchers discovered that adding new tools or changing descriptions confused the agents the most, while removing unnecessary tools actually helped them. They also found that adding "cognitive" features like memory and reflection to the agents helped them adapt better, suggesting that the future of AI isn't just about being smarter, but about being more flexible when the world around them changes.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.