← Latest papers
🤖 machine learning

PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks

This paper introduces PPT-Eval, a benchmark of 120 PowerPoint tasks designed to evaluate computer-use agents, featuring a novel rubric-based framework that captures partial progress and correlates strongly with human judgment to reveal that current frontier models still struggle with complex presentation editing.

Original authors: Apurva Gandhi, Vishwas Suryanarayanan, Raja Hasnain Anwar, Firoz Shaik, Shubhang Desai, Thong Q. Nguyen, Muhammad Taqi Raza, Vishal Chowdhary, Graham Neubig

Published 2026-07-01
📖 4 min read☕ Coffee break read

Original authors: Apurva Gandhi, Vishwas Suryanarayanan, Raja Hasnain Anwar, Firoz Shaik, Shubhang Desai, Thong Q. Nguyen, Muhammad Taqi Raza, Vishal Chowdhary, Graham Neubig

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to use Microsoft PowerPoint. You want to see if the robot can actually look at a slide, understand a request like "add a blue circle here," and then click the right buttons to make it happen.

The paper "PPT-EVAL" is essentially a giant, tricky test designed to see how good these "computer-use" robots are at doing real-world presentation work.

Here is the breakdown of the paper using simple analogies:

1. The Problem: The Robot vs. The Real World

Most previous tests for robots were like giving them a video game controller with a limited set of buttons. The robot could only do what the game allowed (like "add text" or "change color" via a code command).

But in the real world, humans don't use code; they use a mouse, a keyboard, and a screen. They click menus, drag shapes, and use design tools that don't have simple code buttons.

  • The Paper's Claim: Existing tests didn't check if robots could actually use the mouse in PowerPoint. They only checked if robots could use a "cheat code" (API). PPT-EVAL is the first test that forces the robot to use the actual PowerPoint interface (the GUI) just like a human would.

2. The Test: The "PowerPoint Gym"

The researchers built a gym with 120 different exercises (tasks) spread across 12 different PowerPoint files.

  • The Files: They are like real-world decks you might find in a library—some are about medicine, some about history, some about accounting. They have charts, weird layouts, and animations.
  • The Difficulty: The exercises are graded like a video game:
    • Easy: "Change the background color to blue." (Simple)
    • Medium: "Add a new team member to the list and change their font." (A few steps)
    • Hard: "Rearrange this complex diagram, add a new section, and make sure the animation plays correctly." (Requires thinking and many steps).

3. The Scoring: The "Rubric" (The Most Important Part)

This is the paper's biggest innovation. In the past, tests were like a Pass/Fail exam. If the robot got the answer 90% right but missed one tiny detail, it got a zero. That's unfair because the robot did try and did most of the work.

The researchers created a Rubric (a detailed grading sheet), which is like a talent show judge instead of a strict teacher.

  • Partial Credit: If the robot adds the circle but puts it in the wrong spot, the judge gives it points for adding the circle but deducts points for the placement.
  • Penalties: If the robot accidentally deletes a slide or adds a weird animation it wasn't asked for, the judge takes points away.
  • The "Human" Touch: The scoring system uses AI to write a natural language explanation, like: "You got the circle, but it's blocking the text. Good job on the color, but the position needs work."

The Result: This "Talent Show Judge" system was 77% aligned with how actual humans would grade the work. This means the test is fair and accurate.

4. The Results: The Robots Are Still Learning

The researchers put the world's smartest AI robots (like Claude-4.5 and OpenAI's models) through this test.

  • The Score: Even the best robots only got about 45% of the tasks "perfect."
  • The Average: On average, they got about 57% of the points when you count their partial progress.
  • The Comparison: A human expert, given the same test, would score around 80%.
  • The API vs. GUI Gap: Interestingly, when the robots were allowed to use "cheat codes" (APIs) instead of the mouse, they did much better (62% success). This proves that while the robots are smart, they are still clumsy with a mouse and keyboard compared to how they handle code.

Summary

Think of PPT-EVAL as a driving test for robots.

  • Old Tests: Asked the robot, "Can you calculate the speed of the car?" (Code/API).
  • PPT-EVAL: Puts the robot in the driver's seat, hands it the steering wheel, and says, "Drive to the store, park in the spot, and don't hit the curb." (GUI/Mouse).

The paper concludes that while our current robots are getting better, they still struggle with the fine motor skills and visual reasoning required to use PowerPoint like a human. There is still a long way to go before they can fully take over our presentation jobs.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →