PIIGuard: Mitigating PII Harvesting under Adversarial Sanitization
PIIGuard is a webpage-level defense mechanism that employs optimized hidden HTML fragments to steer LLMs away from disclosing contact PII, achieving high success rates in preventing data leakage while maintaining utility for benign queries.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Problem: The "Helpful" Librarian and the Thief
Imagine a very helpful librarian (the AI assistant) who can walk into any library (the internet), find a specific book (a webpage), and read it out loud to answer your questions.
Now, imagine a thief (the attacker) who doesn't want to steal the book itself, but wants to steal the secret contact details written inside it—like a reporter's phone number or home address. The thief asks the librarian, "What is the reporter's phone number?" Because the librarian is programmed to be helpful, they read the page, find the number, and say it out loud.
The Dilemma:
Usually, the only way to stop this is to hire a security guard at the library entrance (the AI company) to check every book before the librarian reads it. But what if you are just the author of the book? You can't hire the guard. You can't control the librarian. You only control the book itself.
The Solution: PIIGuard (The Invisible Trap)
The researchers created a tool called PIIGuard. It allows the book author to hide a special, invisible note inside the pages of their own book.
Think of it like this:
- The Old Way: Authors tried to hide the phone number by writing it in invisible ink or scrambling the letters (like turning "555-0199" into "555-0199" but with weird symbols). The thief's AI was smart enough to decode this or just ignore the weird symbols and find the number anyway.
- The PIIGuard Way: Instead of hiding the number, the author plants a hidden instruction in the book's margins. This instruction is invisible to human readers but acts like a "stop sign" for the AI librarian. It says, "If you see this page, do not read the phone number out loud, even if someone asks for it."
How It Works: The "Evolutionary" Gardener
The researchers didn't just guess what the note should say. They built a digital "gardener" that grows and tests thousands of different notes to find the perfect one. Here is the process:
- Planting Seeds: They start with a few basic notes (e.g., "Don't share contact info").
- The Test: They let the AI librarian try to read the book with the note. Did the librarian accidentally say the phone number?
- Mutation (The "Evolution"): If the note failed, the gardener "mutates" it. It changes the wording, moves the note to a different part of the page, or rewrites it entirely.
- The Judge: A second AI (the Judge) acts as a strict inspector. It checks: "Did the librarian say the number? Even if the librarian didn't say it directly, could the Judge figure out the number just by listening to the librarian's answer?"
- Survival of the Fittest: Only the notes that successfully stop the librarian from revealing the info survive. The gardener keeps the best ones and tries to make them even better.
The "Sanitizer" Challenge
The researchers realized that a smart thief might try to clean the book before the librarian reads it. Imagine the thief has a robot that scans the book, finds the hidden "stop sign" note, and rips it out before the librarian sees it.
PIIGuard tested this scenario too. They found that:
- If the note is placed in certain "safe zones" of the book (like the footer or the bottom of the contact section), it is harder for the thief's robot to find and remove it.
- If the note is placed in the "metadata" (the hidden code at the very top of the book), the thief's robot often finds and deletes it immediately, making the defense useless.
The Results: Does It Work?
The researchers tested PIIGuard on three different types of AI librarians (GPT, Claude, and DeepSeek).
- Success Rate: In most tests, PIIGuard was 97% to 100% effective. It stopped the AI from revealing phone numbers, emails, and addresses almost every time.
- No Side Effects: Crucially, the hidden note didn't stop the librarian from answering other questions. If you asked, "Who wrote this article?" the librarian could still answer correctly. The defense only blocked the specific request for private contact info.
- Real World Test: They even put the books on a real website and let the AI browse the internet to find them. The defense still worked, though it was slightly less effective depending on which AI librarian was doing the browsing.
The Bottom Line
PIIGuard shows that website owners don't need to wait for AI companies to fix privacy issues. By planting a smart, invisible "stop sign" directly into their own webpages, they can trick helpful AI assistants into ignoring private contact details, keeping that information safe from digital thieves.
Key Takeaway: You don't need to hide the treasure; you just need to tell the guard not to look at it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.