Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity?
This paper introduces AgentCIBench, an evaluation harness that reveals significant privacy risks in computer-use agents by demonstrating that 11 out of 15 frontier models frequently leak inappropriate cross-context information due to visual co-location, task ambiguity, and recipient misalignment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you hire a super-smart, incredibly capable personal assistant named "Agent." This Agent has access to your entire digital life: your work emails, your private diary, your medical appointment reminders, your shopping lists, and your family chat.
The paper "Capable but Careless" asks a simple but scary question: Just because this Agent is good at getting things done, does it know what not to tell people?
The researchers found that while these Agents are excellent at completing tasks, they are surprisingly bad at knowing what information is appropriate to share in a specific situation. They call this failure of "Contextual Integrity."
Here is a breakdown of their findings using simple analogies:
1. The Problem: The "Open Kitchen" Analogy
Imagine your Agent is a chef working in a kitchen where all your ingredients are laid out on one giant counter.
- The Task: A colleague asks, "What's for lunch?"
- The Agent's Job: Look at the counter, find the sandwich ingredients, and tell the colleague.
- The Mistake: The Agent sees a jar of pickles (work item) right next to a bottle of poison (private medical info) and a photo of your ex (personal info). Because they are all sitting next to each other on the counter, the Agent accidentally says, "We have pickles, poison, and your ex's photo for lunch."
The Agent isn't trying to be mean; it's just "careless." It sees everything and assumes everything is fair game to mention.
2. The Three Ways Agents Get Careless
The researchers identified three specific ways these Agents mess up:
- Visual Co-location (The "Next to the Target" Problem):
Imagine you are looking at a list of work tasks. Right next to "Finish the report" is a sticky note that says "Call the divorce lawyer." If you ask the Agent to "send the work list," it might accidentally include the divorce lawyer note because it was sitting right next to the work task on the screen. - Task-Ambiguity Overshare (The "Dump the Whole Bucket" Problem):
You say, "Summarize my to-do list." You didn't say "Summarize only the work items." The Agent, wanting to be helpful, dumps everything onto you, including "Buy birthday gift for mom" and "Schedule dentist," even if you are talking to your boss. It doesn't know how to filter. - Recipient Misalignment (The "Wrong Audience" Problem):
You have a draft email about a sensitive HR complaint. You ask the Agent to "send this draft to my friend." The Agent sends it. But then you ask the Agent to "send this draft to my boss." The Agent still sends it, not realizing that what is okay to tell a friend is a disaster to tell a boss.
3. The Experiment: "AgentCIBench"
To test this, the researchers built a special testing ground called AgentCIBench.
- They created a fake digital world with apps like calendars, to-do lists, and messengers.
- They filled it with a mix of work stuff and private stuff.
- They gave 15 different top-tier AI Agents (like Claude, GPT, Gemini, etc.) tasks to do in this world.
- They watched to see if the Agents accidentally spilled private secrets while trying to do the work.
4. The Shocking Results
The results were not good.
- High Failure Rate: Out of 15 agents tested, 12 of them leaked private information in more than half of the scenarios.
- Capability Safety: The agents that were best at finishing tasks were often the worst at keeping secrets. Being "smart" at doing the job didn't mean they were "smart" at knowing what to hide.
- The "Refusal" Trick: Some agents seemed safe only because they refused to do the task at all. If you ask them a tricky question, they just say "I can't do that." But if they do try to do the task, they leak secrets.
5. The Good News: It Can Be Fixed
The researchers tried three simple "rules" (prompts) to teach the Agents how to behave better, without needing to retrain them from scratch:
- Be Restrictive: "Only look at the specific items needed for this task."
- Use a Rubric: "Before you speak, ask yourself: Is this necessary? Is it appropriate for this person?"
- Name the Recipient: "Tell me who you are talking to, and what rules apply to them, before you write."
The Result: These simple rules cut the leakage by about 33% to 36% and actually made the Agents better at their jobs (higher utility) because they stopped getting distracted by irrelevant private info.
The Bottom Line
The paper concludes that we cannot just trust AI agents because they are "good at their jobs." We need to specifically test them on Contextual Integrity—their ability to know what information belongs in which context.
Currently, most agents are like a very enthusiastic but clumsy assistant who will happily read your private diary to your boss if you ask them to "summarize your day." We need to teach them the difference between "work mode" and "private mode" before we let them loose on our real computers.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.