Plover: Steering GUI Agents through Plan-Centric Interaction
Plover is a plan-centric vision-based GUI automation system that enhances transparency and controllability by externalizing task plans as editable artifacts, enabling users to inspect, supervise, and locally correct agent behavior during execution.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where your computer doesn't just sit there waiting for you to click, but actually does things for you. This is the dream of "GUI agents"—smart software helpers that can look at your screen, read what's on it, and click buttons or type text just like a human would. Think of them as digital interns who can navigate websites, fill out forms, or organize files based on a simple sentence you give them, like "Book a flight to Paris."
But here's the catch: these digital interns are still learning. Sometimes they get confused, click the wrong button, or get stuck in a loop, and because they work so fast and silently, you often don't know they've gone off the rails until it's too late. It's like watching someone drive a car blindfolded; if they take a wrong turn, you can't steer them back because you can't see where they are going. The big question researchers are asking is: How do we let these smart helpers do their job without losing control? How do we make sure that when they start to drift, we can gently nudge them back on track without having to restart the whole trip from scratch?
Enter Plover, a new system designed to fix exactly that problem. Instead of letting the computer agent work in secret, Plover acts like a co-pilot that keeps a visible, editable map of the journey. Imagine you're on a road trip with a friend who is driving. In the old way, your friend would just drive, and if they missed a turn, they'd keep driving until they realized they were lost, then maybe panic and start over. Plover changes the game by putting a giant, transparent map on the dashboard that both of you can see. If your friend (the agent) starts heading toward a cliff, you can point at the map and say, "Hey, turn left here," or even draw a line on the map to show the new route. The best part? You don't have to tell them to start the whole trip over; you just fix the part of the map that's wrong, and the car keeps moving forward with the progress you've already made.
The researchers behind Plover, a team from UC Davis and Bosch Research, tested this idea by taking 38 tricky computer tasks that usually cause smart agents to fail. They let the agents try on their own first, and as expected, 26 of them crashed or got stuck. Then, they switched to the Plover mode, where a human could step in, look at the visible plan, and make small, targeted corrections using text, pointing at the screen, or editing the steps directly. The results were promising: by using these "local" fixes, they were able to rescue 23 of those 26 failed tasks. In fact, 17 of them became complete successes, and 6 became partial successes, with the human only needing to step in about twice per task on average.
The study suggests that many of these computer failures aren't permanent disasters; they are just small misunderstandings that can be fixed if we can see what the computer is thinking. The paper argues that the future of helpful AI isn't about making robots that never make mistakes, but about building systems where humans and machines can work together to catch those mistakes early. By making the "plan" visible and editable, Plover turns a frustrating, all-or-nothing experience into a collaborative dance where you can steer the agent back on course without losing all your hard work. It's not a magic bullet that solves every problem—some complex, multi-step messes are still too tangled to fix easily—but it suggests that giving us a better view and a lighter touch is the key to making our digital helpers truly reliable.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.