← Latest papers
💬 NLP

MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents

Original authors: Lawrence Keunho Jang, Andrew Keunwoo Jang, Jing Yu Koh, Ruslan Salakhutdinov

Published 2026-06-16
📖 4 min read☕ Coffee break read

Original authors: Lawrence Keunho Jang, Andrew Keunwoo Jang, Jing Yu Koh, Ruslan Salakhutdinov

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a personal assistant to help you manage your life. You wouldn't give them a brand-new, empty desk with no files, no photos, and no memory of who you are, right? You'd expect them to know your favorite coffee shop, your bank account balance, your upcoming vacation, and your friend's birthday.

The Problem: The "Empty Desk" Test
Currently, the tests we use to see if AI computers are smart enough to be personal assistants are like giving them that empty desk. The AI is asked to perform tasks on a computer that has no history, no logged-in accounts, and no personal data. It's like asking someone to "order your usual dinner" when they have no idea who you are or what you usually order.

The paper argues that this is a huge gap. Real personal assistants need to navigate a digital life that is messy, interconnected, and full of personal history.

The Solution: MYPCBENCH (The "Michael Scott" Simulation)
To fix this, the researchers built a new test called MYPCBENCH. Instead of an empty desk, they created a fully populated digital life for a specific character: Michael Scott from the TV show The Office.

Think of this as building a massive, interactive video game world where:

  • The Character: Michael Scott is the "user."
  • The Data: The computer is pre-loaded with 1,812 bank transactions, 2,398 emails, 679 calendar events, and thousands of chat messages.
  • The Apps: There are 17 different "fake" websites (like a fake bank, a fake airline, a fake food delivery app) that are all connected. If Michael books a flight on the fake airline, a charge automatically appears on his fake bank statement, and a confirmation email lands in his fake inbox.
  • The Goal: The AI has to act as Michael's assistant, navigating these apps to solve real-life problems using his specific history.

The Test: 184 Real-World Challenges
The researchers created 184 specific tasks based on real requests people actually make to AI assistants. These tasks range from simple to incredibly complex:

  • Simple: "Send $100 to Pam via Zelle." (The AI has to check if Pam is in the contacts first).
  • Complex: "I have trips to Jamaica and Barbados booked. Can I afford both based on my current credit card balance?" (The AI has to look at the bank app, the travel app, and do some math).
  • Pattern Finding: "What do I usually tip on food delivery?" (The AI has to scan hundreds of past orders to find a pattern).

The Results: The AI Struggles with "Real Life"
The researchers tested six of the smartest AI models available (from companies like Anthropic, OpenAI, and others) on this test. Here is what happened:

  1. The Best Performer: The top model (Claude Opus 4.6) managed to solve about 55% of the tasks perfectly. That sounds good, but remember, this is the best AI available. It still failed nearly half the time.
  2. The "Multi-App" Problem: The AI got significantly worse as tasks required using more than one app. When a task needed the AI to jump between 7 or more different apps (like checking email, then a calendar, then a bank, then a travel site), the success rate for most models dropped to 0%.
  3. The "Long Memory" Problem: The AI struggled to keep track of long chains of actions. If a task required 50 steps, the AI often got lost or gave up.
  4. Different Failure Styles:
    • Some AIs (like GPT) tended to give up too early, saying "I'm done" before actually finishing the job.
    • Some AIs (like Qwen) started making things up, inventing facts about Michael Scott that weren't in the data.
    • The best AI (Claude) sometimes took shortcuts by using the computer's command line (like a hacker) instead of clicking through the menus like a normal person.

The Bottom Line
The paper concludes that while AI is getting better at using computers, it is still not ready to be a true "personal" assistant. It can handle a clean, simple task, but it falls apart when it has to navigate a messy, interconnected, real-world digital life with personal history.

The researchers released their test environment (the "Michael Scott" computer) to the public so other scientists can try to build better assistants that can actually handle the complexity of a real human's digital life.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →