What Your Posts Reveal: A Benchmark and Agentic Framework for User-Level Privacy Leakage on Social Media
This paper introduces SopriBench, a synthetic benchmark for user-level multimodal privacy leakage, and Argus, a training-free agentic framework that leverages abductive reasoning to achieve a 25% improvement in cumulative cross-post privacy inference over existing baselines.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: The "Jigsaw Puzzle" of Privacy
Imagine you are playing a game of hide-and-seek, but instead of hiding a person, you are hiding your home address.
Most people think that if they don't explicitly say, "I live at 123 Maple Street," they are safe. But this paper argues that privacy leaks are like a jigsaw puzzle.
- Post 1: You post a photo of your morning coffee. (Harmless)
- Post 2: You post a picture of your commute, showing a specific subway station. (Harmless)
- Post 3: You post a photo of your lunch, showing a unique restaurant sign. (Harmless)
Individually, these posts are just boring daily updates. But if someone puts them together, they can figure out exactly where you live, where you work, and your daily routine. This is called cumulative leakage. The danger isn't in one big secret; it's in hundreds of tiny, harmless clues that add up to a complete picture.
The Problem: We Didn't Have a Good Test
The researchers noticed two major problems in how we study this:
- No Good "Exam" (Benchmark): Previous tests only looked at single posts (e.g., "Does this one photo reveal a name?"). They didn't test if an AI could solve the whole puzzle by connecting clues across 50 different posts.
- Wrong Grading System: Old tests gave a "Pass/Fail" grade. If an AI guessed your city, it got a point. If it guessed your exact street, it also got a point. The researchers say this is unfair. Guessing your city is a small leak; guessing your exact apartment number is a massive, dangerous leak. We need a way to measure how bad the leak is, not just if it happened.
The Solution: SopriBench (The Test)
To fix this, the team built SopriBench.
- What is it? A fake, synthetic social media dataset. They didn't use real people's data (to protect privacy). Instead, they studied thousands of real posts from apps like Rednote and Instagram to understand how people accidentally leak info. Then, they used that knowledge to create 50 fake user profiles with 500 posts and 1,569 images.
- The "Privacy Exposure Score" (PES): This is their new grading system.
- If an AI guesses your Country, it gets a low score.
- If it guesses your City, it gets a medium score.
- If it guesses your Exact Home Address, it gets a high score.
- Analogy: It's like a weather report. Knowing it's "raining" is okay. Knowing it's "raining specifically on your roof at 3 PM" is much more specific and potentially more useful to a burglar.
The New Detective: Argus
The researchers also built a new AI detective called Argus.
- How other AIs work: Most AIs are like a photocopier. You show them a post, and they immediately say, "I see a coffee cup, so you like coffee." They look at each post in isolation and make a quick guess.
- How Argus works: Argus is like a private investigator or a detective.
- The Skim: It looks at all 50 posts quickly and writes down every tiny clue (a ticket, a background sign, a timestamp).
- The Hypothesis: It forms a theory: "Maybe this person lives in Shanghai?"
- The Investigation: It doesn't just guess. It goes back to other posts to check. "Wait, Post 12 showed a subway map. Does that match Shanghai?" It uses tools (like a map search or reading text in an image) to verify its theories.
- The Verdict: Only when the evidence from multiple posts lines up does it write down the final answer.
The Result: Argus was much better at solving the "puzzle" than the other AIs. It improved the accuracy of finding private info by 25%, especially when the clues were scattered across different posts.
Why This Matters (According to the Paper)
The paper concludes that we need to stop looking at social media privacy as "one post at a time."
- The Takeaway: Just because you don't post your address doesn't mean you are safe. If you post enough small, harmless details over time, a smart system (or a human) can connect the dots to find your deepest secrets.
- The Warning: The paper shows that current AI tools are getting very good at being these "connectors." If we don't develop better ways to measure and stop this, our digital footprints are much more revealing than we think.
In short: The paper built a better test (SopriBench) and a smarter detective (Argus) to prove that privacy leaks are a cumulative game of connect-the-dots, and the picture they draw is often much more detailed than we realize.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.