Benchmarking Living-Screen-Native GUI Agents on Short-Video Platforms
This paper introduces LivingScreen, the first benchmark for "Living-Screen-Native" GUI agents that operate on dynamic short-video platforms, revealing that current frontier models fail to match human performance due to an inability to effectively control observation timing amidst continuously changing content.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to use a smartphone. Most current tests for these robots are like giving them a frozen photograph of a screen. The robot looks at the picture, decides what to click, and then the picture changes. It's a bit like playing a game of "Simon Says" where the board never moves unless you touch it.
But real life isn't frozen. Think of TikTok, Instagram Reels, or YouTube Shorts. The videos play automatically, comments scroll by, and new content loads while you are still watching. The screen is "alive."
This paper introduces a new way to test robots (called GUI Agents) on these "living" screens. Here is the breakdown in simple terms:
1. The Problem: The "Frozen Screen" vs. The "Living Screen"
Current AI models are great at looking at a static image or a pre-recorded video file. But they struggle when the screen is constantly changing on its own.
- The Old Way: The AI gets a snapshot, thinks, clicks, and then gets a new snapshot.
- The New Way (Living-Screen-Native): The screen keeps playing. The AI has to decide: "Should I watch this whole video? Should I just glance at the first 3 seconds? Should I skip this one entirely?"
The authors call this "Living-Screen-Native" because the agent has to navigate a screen that evolves in real-time, just like a human does.
2. The Solution: "LIVINGSCREEN" (The New Test)
To test this, the researchers built a special playground called LIVINGSCREEN.
- The Environment: It's a realistic fake version of a short-video app inside a web browser. It plays real videos, shows comments, and has buttons for "like," "share," and "follow."
- The Rules: The AI can't just "download" the video file. It has to interact with the screen exactly like a human: swiping up, clicking buttons, and deciding how long to watch a clip.
- The Tasks: The test has three levels of difficulty:
- Level 1 (Basic): Can you click the "like" button or scroll down?
- Level 2 (Understanding): Can you watch a few videos, read the comments, and figure out if two videos tell the same story?
- Level 3 (The Boss Fight): Can you act like a human user? For example, "Find videos that break the rules and report them," or "Find videos about cooking and save the best ones."
3. The Results: The Robots Are Clumsy
The researchers tested the smartest AI models available today on this new test. The results were surprising:
- Humans Win: Humans are much better at balancing speed and accuracy. We know exactly how long to watch a video to get the answer.
- AI Struggles: Even the best AI models failed to match human performance. They were either too slow (watching too much) or too hasty (missing important details).
4. The Big Discovery: "Over-Observing" and "Under-Observing"
The paper found that the main reason AI fails isn't that it can't see or click; it's that it doesn't know how to watch. The authors call this "Over- and Under-Observation."
Think of it like a student taking a test:
- Under-Observation: The student glances at a question, guesses the answer immediately, and moves on, missing the crucial detail in the middle of the paragraph.
- Over-Observation: The student reads the same paragraph five times, even after they already understood it, wasting time and energy.
The AI's Problem:
- Some AIs glance too little. They skip videos they should have watched, leading to wrong answers.
- Some AIs watch too much. They stare at irrelevant videos for minutes, wasting time, without actually getting smarter about the task.
- The Human Difference: Humans are "smart scanners." We do a quick 2-second "glance" to see if a video is relevant. If it is, we watch longer. If not, we swipe away. The AI models mostly lack this "quick glance" filter. They tend to either ignore a video completely or binge-watch it.
5. Can We Just Tell Them to Do Better?
The researchers tried giving the AI models simple instructions, like "Watch less" or "Watch more," or even telling them to "Copy how humans watch."
- The Result: It didn't work well. Telling the AI to "watch less" made it skip too many important videos. Telling it to "watch more" made it waste time.
- The Conclusion: This isn't just a matter of giving better instructions. The AI models genuinely lack the capability to control their own attention. They don't know how to decide when to look and when to stop looking.
Summary
This paper says that to build truly helpful AI assistants for apps like TikTok or YouTube, we need to stop testing them on frozen screens. We need to teach them how to manage their attention in a world that never stops moving. Currently, the best AI models are like a person who either never looks at the menu or reads every single word on it for an hour—they just haven't learned how to "glance" effectively yet.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.