How Benchmarks Mis-Score Computer-Use Agents
This paper critiques the reliability of current computer-use agent benchmarks by identifying systemic flaws in task construction, observation, and scoring that lead to erroneous failure verdicts, and proposes a diagnostic framework and design rules to improve evaluation accuracy.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are watching a high-stakes cooking competition on TV. The judges are supposed to taste the dishes and decide who wins. But what if the judges are using a broken taste test? What if the recipe they are following is written in disappearing ink, or the oven they are using is actually a toy that doesn't heat up? In the world of artificial intelligence, specifically with "Computer-Use Agents" (smart bots that can click, type, and browse the web just like humans), we have a similar problem. These bots are being tested to see if they can do real jobs, like filing taxes or booking flights. However, the "scorecards" we use to grade them are often glitchy. They might give a failing grade to a bot that actually did a great job, or they might miss the fact that the bot failed because the test itself was broken. This paper is like a detective story where the authors investigate why these scorecards are lying to us.
The paper, titled "How Benchmarks Mis-Score Computer-Use Agents," argues that the way we currently test these AI bots is deeply flawed. The authors found that when a bot gets a "FAIL" score, there is a surprisingly high chance that the score is actually wrong. In their investigation of 150 specific examples where bots were marked as failures, they discovered that 15.3% of those "FAIL" verdicts were mistakes.
Here is how the mistakes happened:
- The Judge was Too Picky (10.7%): Sometimes the bot did the right thing, but the automated checker was too rigid. It was like a judge failing a chef because they used a slightly different knife than the one on the recipe card, even though the food tasted perfect.
- The Test was Broken (4.7%): Sometimes the bot couldn't finish the task because the test environment was broken. It was like asking a chef to bake a cake in an oven that had no power. The bot failed, but not because it was bad at cooking; the kitchen was broken.
The authors also looked at the bots that actually failed. They found that the main reasons weren't usually that the bot clicked the wrong button (which is like a clumsy hand). Instead, the bots mostly failed because they got stuck in a loop of bad planning or they didn't realize their actions weren't working (like a chef who keeps stirring a pot that is already empty).
The paper suggests that we can't just look at a single number, like a "pass rate," to know if an AI is good. That number hides too many secrets. Instead, we need to build better tests that are harder to break, use smarter judges that understand different ways to solve a problem, and provide full video recordings of the tests so we can see exactly where things went wrong. The goal is to stop giving bots a failing grade for things that aren't their fault, and to help us understand what they actually need to learn to become truly helpful helpers.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.