Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?
The paper introduces Skill-Use, a benchmark evaluating whether LLM agents can independently recognize and correctly apply structured skills in isolated environments, revealing that reliable skill use remains a significant challenge heavily dependent on the specific agent harness rather than being an inherent model capability.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a super-smart, hyper-creative intern who can write code, analyze data, and draft emails. You give them a massive library of "Skill Books." Each book is a recipe for a specific job, like "How to fix a broken server" or "How to write a legal contract." The catch? You don't hand them the whole book at once. Instead, you only show them the title and a tiny blurb on the back cover. If they think the book is relevant, they have to ask for it, find the full instructions, and then follow those instructions exactly without skipping steps or doing anything forbidden. This is the world of AI "agents" trying to use "skills." For a while, people hoped that if you just gave these AI interns the right books, they would instantly become perfect workers. But nobody really knew if the AI could actually figure out which book to grab, or if it would just guess and hope for the best.
A team of researchers from Tencent Hunyuan and several top universities decided to put this idea to the test. They built a giant playground called SKILL-USE, a benchmark designed to see if AI agents can truly "read the room" and use these skill books correctly. They didn't just ask, "Did the AI finish the task?" They asked three much harder questions: Trigger (Did the AI even know to grab the right book?), Compliance (Did it follow the recipe step-by-step?), and Boundary (Did it avoid doing the forbidden things listed in the book?). They tested 177 different real-world tasks, from fixing software bugs to writing business reports, using 79 real skill documents and eight of the smartest AI models available today.
The results were a bit of a reality check. The researchers found that even the best AI setups only managed to get a "Skill-Use" score of about 0.613 (on a scale of 0 to 1). That means reliable skill use is still out of reach. The biggest problem wasn't that the AI couldn't follow the rules once it had the book; it was that the AI often didn't even know to pick the book up in the first place. It's like having a genius chef who can follow a recipe perfectly, but who keeps forgetting to open the fridge to get the ingredients. Furthermore, the researchers discovered that an AI's performance wasn't just about how "smart" the model was; it depended heavily on the "harness" or the software wrapper around it. Changing the wrapper was like changing the kitchen layout: sometimes the same chef became a star, and other times they stumbled.
In short, the paper suggests that giving AI agents a library of skills doesn't automatically make them better workers. The gap between "having a skill" and "using a skill" is huge. The AI often fails to recognize when a skill applies, or it gets distracted by the wrong skills when there are too many choices. While the AI is getting better at avoiding forbidden moves, it still struggles to follow complex procedures faithfully. The authors conclude that we can't just assume these agents will magically learn to use their tools; we need to build systems that help them recognize when to use a skill and how to stick to the plan, because right now, that's the hardest part of the puzzle.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.