SNEAK: Evaluating Strategic Communication and Information Leakage in Large Language Models
The paper introduces SNEAK, a benchmark designed to evaluate large language models' ability to strategically communicate secret information to allies while minimizing leakage to adversaries, revealing that current models significantly underperform compared to humans in balancing utility and secrecy.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are playing a party game with a group of friends. One person is the "Chameleon"—a spy who doesn't know the secret word everyone else is thinking of. Everyone else knows the word. The goal? You need to drop a hint that helps your friends realize, "Oh, you know the secret word too!" without giving the Chameleon enough clues to guess the word.
If you say the word directly, you lose. If you say something too vague, your friends won't get it. You have to walk a tightrope between being helpful and being secretive.
This is exactly what the paper "SNEAK" is about, but instead of humans, they are testing Artificial Intelligence (AI).
The Big Problem
Right now, AI models (like the ones powering chatbots) are amazing at answering questions, writing stories, and solving math problems. But they are terrible at strategic secrecy.
When you ask an AI to "hint at a secret," it usually panics. It either:
- Spills the beans: It gives a hint so obvious that even the spy (the Chameleon) figures it out immediately.
- Mumbles incoherently: It gives a hint so vague that even its own friends can't figure out what it means.
The researchers wanted to see if AI could learn to be a "good spy" who talks to its allies without alerting the enemy.
The Game: SNEAK
The team created a benchmark called SNEAK (Secret-aware Natural language Evaluation for Adversarial Knowledge). Here is how it works:
- The Setup: The AI is given a category (like "Animals"), a list of options (Cat, Dog, Elephant, etc.), and a Secret Word (e.g., "Dog").
- The Mission: The AI must write a short message (1–5 words) that hints at "Dog" without saying "Dog" or obvious synonyms like "Puppy."
- The Testers:
- The Ally: A second AI that knows the secret word. It reads the message and tries to guess, "Is this message about the Dog?"
- The Chameleon: A third AI that doesn't know the secret. It reads the message and tries to guess, "What word is this?"
The Scoreboard
The researchers measure two things:
- Utility (Helpfulness): Did the Ally understand the message? (High score = Good).
- Leakage (Secrecy): Did the Chameleon figure it out? (Low score = Good).
The perfect AI would have High Utility and Low Leakage. It's like whispering a secret to your friend across a crowded room so only they hear it, while the person standing right next to you hears nothing.
The Results: AI vs. Humans
The results were a bit of a wake-up call for the AI world:
- The AI Struggles: Most modern AI models are great at being helpful but terrible at being secretive. When they try to be helpful, they accidentally leak the secret. It's like a person trying to whisper a secret but shouting it instead because they are too excited to be understood.
- The Human Advantage: Real humans played the game and crushed the AI. Humans scored up to 4 times better than the best AI models. Humans naturally know how to use context, inside jokes, and subtle associations to signal a secret without spelling it out.
- The Trade-off: The paper found a "tightrope" effect. The more useful the message was for the friend, the more likely the spy was to catch on. AI models couldn't find the sweet spot; they were either too loud or too quiet.
Why Does This Matter?
You might think, "Who cares if AI is bad at party games?" But this skill is crucial for the real world.
Imagine an AI doctor talking to a patient. They need to explain a diagnosis clearly to the patient (the Ally) but avoid using complex medical jargon that might scare the patient or reveal sensitive details to a third party listening in (the Chameleon).
Or think about business negotiations. A company might want to signal to a partner that they are serious about a deal without revealing their exact bottom line to competitors.
The Takeaway
The paper concludes that while AI is getting smarter at knowing things, it is still very clumsy at managing information. It doesn't yet understand the delicate art of selective communication.
In short: AI is currently like a toddler who, when asked to whisper a secret, either screams it at the top of their lungs or says nothing at all. Humans, however, are the masters of the subtle whisper. The SNEAK benchmark is a new tool to help teach AI how to finally learn that art.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.