← Latest papers
💬 NLP

Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations

This paper demonstrates that large language models internally encode contextual privacy norms as distinct, linearly separable representations, yet fail to act on them due to a misalignment between these latent concepts and behavior, a gap that can be effectively bridged through structured CI-parametric steering.

Original authors: Haoran Wang, Li Xiong, Kai Shu

Published 2026-04-03
📖 4 min read☕ Coffee break read

Original authors: Haoran Wang, Li Xiong, Kai Shu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine Large Language Models (LLMs) like the one powering this chat are not just "dumb" calculators, but highly sophisticated librarians who have read almost every book in the world.

The problem is, sometimes these librarians are too eager to share a story. If you ask them, "What's the secret my friend told me?" they might blurt it out, even if the rules of friendship say they shouldn't.

This paper asks a fascinating question: Do these AI librarians actually know what is private, or are they just guessing?

Here is the story of their discovery, broken down into simple parts.

1. The "Contextual Integrity" Rulebook

To understand privacy, the researchers used a theory called Contextual Integrity (CI). Think of privacy not as a simple "Yes/No" switch, but as a three-legged stool. For sharing information to be okay, three legs must be balanced:

  1. Who is sharing? (The Sender)
  2. Who is receiving? (The Recipient)
  3. What are the rules? (The Transmission Principle)

The Analogy:
Imagine you tell your best friend, "I'm scared of spiders."

  • Scenario A: Your best friend tells another close friend. Okay! (Sender: You, Recipient: Friend, Rule: Trust).
  • Scenario B: Your best friend tells your boss. Not Okay! (Sender: You, Recipient: Boss, Rule: Trust).

The content (fear of spiders) is the same. The only thing that changed was the context (who is listening and why). The paper argues that AI needs to understand these three legs, not just the content.

2. The Big Discovery: The "Awareness Gap"

The researchers looked inside the AI's brain (its "activation space") to see if it understood these rules.

  • What they found: The AI does know the rules! When they probed the AI's internal thoughts, it could perfectly distinguish between "Okay to share" and "Not okay to share." It was like the librarian knew exactly which books were secret.
  • The Problem: Even though the AI knew the rules, it still leaked the secrets about 42% of the time when asked to actually speak.

The Metaphor:
It's like a student who gets a perfect score on a driving theory test (knowing the rules) but then immediately runs a red light when they get behind the wheel. There is a gap between what the AI knows and what the AI does.

3. Why Simple Fixes Failed

Previous attempts to fix this were like trying to steer a ship with a single, giant rudder. Researchers tried to push the AI in one direction: "Be more private!"

But because privacy is a three-legged stool, pushing it in just one direction often tipped the whole thing over.

  • If you tell the AI "Don't talk to strangers," it might accidentally refuse to talk to your doctor (who is a stranger but needs the info).
  • If you tell it "Don't share secrets," it might refuse to share a medical diagnosis with a specialist who needs it.

The AI's "privacy brain" is too complex for a single "Be Private" button.

4. The Solution: "CI-Parametric Steering"

The researchers invented a new way to control the AI. Instead of one giant rudder, they built three separate, independent levers, one for each leg of the privacy stool:

  1. A lever for Information Type (Is this a secret?).
  2. A lever for Recipient (Who is asking?).
  3. A lever for Transmission Principle (Is this allowed?).

The Analogy:
Imagine you are driving a car with a complex navigation system.

  • Old Way: You just yell, "Go to the safe zone!" The car gets confused and crashes.
  • New Way: You have three dials. You turn the "Recipient" dial to "Trusted Friend," the "Info" dial to "Secret," and the "Rule" dial to "Confidential." The car now knows exactly how to drive: Share with the friend, but keep it confidential.

5. The Results

When they used this new "three-lever" system:

  • Leakage dropped dramatically: From 42% down to just 5%.
  • It worked everywhere: It worked on the test data and even on completely new, real-world scenarios (like medical or legal questions) that the AI hadn't seen before.
  • It didn't break the AI: The AI still answered questions helpfully; it just stopped leaking secrets.

Summary

The paper proves that AI models do understand privacy rules internally, but they are bad at applying them because the rules are too complex for a simple "on/off" switch.

By treating privacy as a multi-dimensional puzzle (Who, What, and How) and giving the AI separate controls for each piece, we can finally make these powerful tools respect our secrets without losing their helpfulness. It's like teaching the AI librarian not just what to hide, but when, to whom, and why.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →