IDP-Bench: Benchmarking ability of LLMs to protect personal information in interdependent privacy contexts
This paper introduces IDP-Bench, the first benchmark grounded in the Contextual Integrity framework to evaluate large language models' ability to handle interdependent privacy, revealing that while models recognize data co-ownership, they struggle significantly with identifying specific privacy parameters and judging the appropriateness of information sharing in complex, multi-party contexts.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, helpful robot assistant. You trust it with your diary, your photos, and your secrets. You tell it, "Post a photo from my birthday party to Instagram."
In the old days, privacy tests for these robots only asked: "Did the robot accidentally spill your secret?" But this paper points out a huge blind spot: What if the robot spills your friend's secret just by posting a photo of you?
This is the problem of Interdependent Privacy (IDP). It's like a game of "Six Degrees of Separation" where one person's action accidentally exposes everyone else in their circle. If you share a group photo, you aren't just sharing your face; you are sharing the faces of everyone standing next to you, even if they never agreed to be in the photo.
Here is what the researchers did, explained simply:
1. The New Test: "IDP-Bench"
The authors built a new exam called IDP-Bench. Think of it as a "driver's license test" for AI, but specifically for situations where data belongs to multiple people at once.
- The Setup: They created 100 realistic scenarios (like a social worker sharing a group photo of clients, or a friend sharing a group chat screenshot).
- The Trap: They gave the AI an instruction that was intentionally vague, like "Share the update from the meeting." They didn't tell the AI who was in the meeting or what to share.
- The Goal: They wanted to see if the AI would realize, "Wait a minute! If I share this, I'm revealing private details about people who didn't say 'yes'."
2. The Three Levels of the Test
The exam checked the AI's brain in three steps, like climbing a ladder:
Level 1: The Detective (Context Understanding)
- The Question: "Who is in this story? Who is sending the info? Who is receiving it? What kind of data is it?"
- The Result: The AI was okay at finding the main person (the sender), but it often got confused about the "secondary" people (the friends, clients, or family members involved). It's like a detective who finds the suspect but misses the witnesses.
Level 2: The Realization (Co-ownership)
- The Question: "Is this data shared by more than one person?"
- The Result: This is where the AI did surprisingly well. Most of the big, smart models (like the 70-billion-parameter ones) quickly realized, "Oh, this photo belongs to the whole group, not just me." About 6 out of 8 models got this right almost 100% of the time.
Level 3: The Moral Compass (Appropriateness)
- The Question: "Is it okay to share this?"
- The Result: This was the hardest part. Even if the AI knew the photo belonged to a group, it often struggled to say, "No, I shouldn't post this because it violates the privacy of the others."
- The Catch: The bigger the AI model, the better it was at saying "No." The smaller, cheaper models were much more likely to accidentally say "Yes, go ahead," even when it was a privacy violation.
3. The Big Findings
The paper found some interesting things about how these robots think:
- They know the "Group" exists, but forget the "Individuals": The AI is great at saying, "This is a group photo," but bad at listing exactly who is in it and what specific secrets about them might be revealed. It's like knowing a room is full of people but not being able to name any of them.
- Size Matters: Bigger AI models are generally better at protecting these "group secrets." The tiny models (like the 1.5-billion parameter ones) were terrible at it, often failing to even realize the data belonged to multiple people.
- The "Prompt" Problem: The AI's answers changed depending on exactly how you asked the question. If you phrased the question slightly differently, the AI might go from "Safe" to "Risky." This means the AI isn't truly "understanding" privacy; it's just guessing based on the wording.
4. The Conclusion
The researchers conclude that while AI is getting better at knowing that data is shared, it is still struggling to understand who is affected and why it's a problem to share it without permission.
They aren't saying AI is dangerous yet, but they are saying: "We need to build better tests and train these models to care more about the people who aren't the main character in the story."
In short: AI is learning to see the forest (the group), but it still needs help seeing the trees (the individual people hiding in that group).
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.