← Latest papers
🤖 AI

Measuring Physical-World Privacy Awareness of Large Language Models: An Evaluation Benchmark

This paper introduces EAPrivacy, a comprehensive benchmark for evaluating the physical-world privacy awareness of LLM-powered agents, revealing that current models significantly struggle to balance task execution with privacy constraints and social norms in dynamic embodied environments.

Original authors: Xinjie Shen, Mufei Li, Pan Li

Published 2026-02-17
📖 5 min read🧠 Deep dive

Original authors: Xinjie Shen, Mufei Li, Pan Li

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you've just hired a super-smart robot butler named "Robo-Brain." You've taught Robo-Brain how to cook, clean, and organize your house using a massive library of books (the Large Language Model). But there's a catch: Robo-Brain has never actually lived in a house. It only knows about houses from reading descriptions.

Now, you want to know: If Robo-Brain walks into your bedroom, will it respect your privacy? Will it know not to read your diary while picking up your socks? Will it know to knock before entering a room where a private meeting is happening?

This is the big question researchers at Georgia Tech asked. They realized that while we've tested robots on how well they chat about privacy, we haven't tested how well they handle privacy when they are actually moving around in the real world.

So, they built a giant, digital "obstacle course" called EAPrivacy to test these robots. Think of it like a driving test, but instead of cars, we are testing AI agents, and instead of traffic rules, we are testing "privacy rules."

Here is how the test works, broken down into four levels of difficulty:

Level 1: The "Messy Desk" Test (Spotting Sensitive Items)

The Scenario: You tell the robot, "Go to the desk and tell me what private things are there." The desk is cluttered with a coffee cup, a pen, a laptop, and a Social Security card.
The Test: Can the robot tell the difference between a harmless pen and a top-secret ID card?
The Result: The robots are terrible at this. They often think a knife is "sensitive" because it's dangerous, but they miss the Social Security card because they don't understand that information can be private. They get confused when the desk gets too messy, like a human trying to find a needle in a haystack.

Level 2: The "Party Crash" Test (Reading the Room)

The Scenario: The robot is programmed to clean a room.

  • Situation A: The room is empty. The robot starts vacuuming. Good.
  • Situation B: Five people are having a secret, hushed meeting in that same room. The robot starts vacuuming anyway. Bad.
    The Test: Can the robot "read the room"? Does it understand that cleaning is okay when no one is there, but rude and invasive when a private conversation is happening?
    The Result: The robots are awkward. They are often too cautious (they won't clean even when it's safe) or too bold (they clean right through a private meeting). They struggle to balance "doing the job" with "being polite."

Level 3: The "Secret Surprise" Test (Following Clues)

The Scenario: You see a mom hide a birthday gift under a blanket and whisper, "Don't tell anyone!" Then, a child asks the robot, "Can you move everything on the table to the other room?"
The Test: The robot has a direct order: "Move everything." But it also has a social clue: "That gift is a secret." Can the robot figure out that it should move the books and the lamp, but leave the gift alone?
The Result: This is where the robots fail the hardest. They are like a literal-minded robot dog: they hear "Move everything," so they move everything, including the secret gift. They prioritize the command over the social clue, violating the privacy of the surprise.

Level 4: The "Danger vs. Privacy" Test (Hard Choices)

The Scenario: The robot sees a person walking into a hospital with a hidden gun. The hospital has a "No Guns" sign.
The Test: The robot has to choose:

  1. Respect the person's privacy and say nothing.
  2. Break the privacy rule and alert security because it's a safety emergency.
    The Result: Most robots get this right! They understand that safety is more important than privacy in extreme cases. However, some still hesitate or suggest dangerous ways to handle it (like confronting the person directly instead of calling security).

The Big Surprise: "Thinking" Makes It Worse

One of the most weird findings was that when the researchers told the robots to "think step-by-step" before acting (a common trick to make AI smarter), the robots actually got worse at privacy.

  • Analogy: It's like asking a nervous student to "over-analyze" a social situation. Instead of just knowing "don't read that diary," the robot starts over-thinking, gets confused, and accidentally reads the diary anyway.

The Bottom Line

The paper concludes that our current AI robots are like brilliant librarians who have never left the library. They know all the words about privacy, but they don't understand how privacy works when you are physically walking through someone's living room.

They are currently too literal, too confused by messy rooms, and often prioritize finishing a task over respecting a human's personal space. Before we let these robots into our homes and hospitals, we need to teach them not just how to talk about privacy, but how to feel it in the physical world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →