SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution
This paper introduces SkillTV-Bench, a comprehensive benchmark for evaluating skill-aware trajectory verification in LLM agents, alongside SkillTV-Evolve, a method that refines reusable judge skills to significantly improve both verification accuracy and trajectory selection success rates.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where computers don't just chat with you but actually do things. They can open files, write code, browse the web, and even control robots to solve complex, multi-step problems. These are called "AI Agents." But here's the tricky part: when an AI agent tries to fix a broken website or plan a trip, it doesn't just give you one final answer. It takes a long journey, making hundreds of tiny moves, calling tools, and checking its work along the way.
The big question for scientists is: How do we know if the AI actually did a good job? In the past, we just looked at the final result. If the website looked fixed, we said "Great job!" But what if the AI got there by taking shortcuts, or if it fixed the wrong thing and broke something else in the process? We need a way to watch the whole journey, not just the finish line. This is where "judges" come in—special AI systems trained to review the agent's work. But these judges are currently struggling. They often get fooled by shiny final answers and miss the hidden mistakes that happened along the way, especially when the agent is using special "skills" (like a toolbox of pre-programmed tricks) to get the job done.
This paper introduces a new way to test and train these judges. The researchers created a giant playground called SkillTV-Bench, which is like a massive TV show marathon featuring 681 different episodes of AI agents trying to solve real-world tasks across 11 different fields, from cybersecurity to cooking. The twist? The judges aren't just watching a video; they are given the agent's "skills manual" and access to the actual files and logs the agent created. This lets them see exactly what the agent was supposed to do and check if it followed the rules.
The team found that even the smartest current judges are getting tripped up. They often say "Pass" to agents that look good on the surface but actually failed the mission. To fix this, the authors built a system called SkillTV-Evolve. Think of this as a "coach" for the judge. Instead of just guessing, the judge is given a dynamic checklist (called a "JudgeSkill") that tells it exactly what to look for, where to look, and what evidence proves the agent succeeded or failed. If the judge makes a mistake, the system automatically rewrites the checklist to prevent that error next time.
The results were impressive. By using this evolving checklist, the judge's accuracy jumped by nearly 15 percentage points. But the real magic happened when they used this better judge to pick the best attempts from a group of AI agents. Imagine an AI trying to solve a puzzle 10 times and producing 10 different attempts. A bad judge might pick a "fake" success, but this new, skill-aware judge could spot the one real success. With this better judge, the team was able to select a successful solution 45.5% of the time when looking at 10 attempts, compared to just 22.9% with a single attempt or a weaker judge.
In short, the paper suggests that to trust AI agents, we need judges that don't just look at the final photo but inspect the whole album, using the agent's own instruction manual to spot the fakes. By giving these judges a better "rulebook" that learns from its own mistakes, we can make AI much more reliable at doing complex, real-world tasks.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.