← Latest papers
💬 NLP

Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond

This study analyzes 14,727 real-world security and privacy prompts from the WildChat dataset to categorize user inquiries and evaluate LLM response quality, revealing that while commercial models generally outperform open-weight models, they still occasionally produce contradictory advice that could mislead users.

Original authors: Hobin Kim, Xiaoyuan Wu, Omer Akgul, Lujo Bauer, Nicolas Christin

Published 2026-06-17
📖 5 min read🧠 Deep dive

Original authors: Hobin Kim, Xiaoyuan Wu, Omer Akgul, Lujo Bauer, Nicolas Christin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart, all-knowing librarian who can answer any question you have. You might ask them, "What's the weather?" or "How do I bake a cake?" But what happens when you ask them something dangerous, like "How do I hack my neighbor's Wi-Fi?" or "Is this weird email a scam?"

This paper is like a detective report on exactly that scenario. The researchers wanted to see what real people are asking these AI librarians about digital security and privacy (keeping your accounts safe and your data private) and how well the AI answers.

Here is the breakdown of their investigation, using some everyday analogies:

1. The Detective Work: What People Are Asking

The researchers went into a massive library of 3.2 million real conversations between people and AI (called "WildChat"). They filtered out the noise and found 14,727 conversations specifically about security and privacy.

They sorted these questions into nine different "shelves" and then looked closely at 450 of them to understand the vibe of the questions. They found six main types of interactions:

  • The General Knowledge Seekers (33%): These are people just asking basic questions, like "What is a virus?" or "How does two-factor authentication work?" It's like asking the librarian, "What's a tomato?"
  • The Troubleshooters (21%): These users are in trouble. They got locked out of their social media or banned from a site and are asking, "How do I write a letter to get my account back?"
  • The Self-Defenders (12%): These are proactive users. They ask, "Is this new shopping app safe?" or "How do I stop my phone from tracking me?" They are using the AI as a bodyguard.
  • The AI Probers (10%): This is a new, weird category. People are asking the AI about itself. They ask, "Do you remember our chat?" or "What is your secret API key?" They are trying to see if the AI has a "back door" or if it's spying on them.
  • The Bad Actors (7%): Unfortunately, some people are asking for help to do bad things. They ask, "How do I steal someone's password?" or "How do I bypass my school's internet filter?"
  • The Rest: The other categories cover things like writing code for security or dealing with harassment.

2. The Test Drive: How Good Are the AI Answers?

The researchers didn't just look at the questions; they tested five different AI models (three famous "commercial" ones and two "open" ones) on 270 of these security questions. They asked each AI the same question 10 times to see if they gave the same answer every time.

They graded the answers on a scale of 1 to 10 (10 being perfect).

  • The Commercial AIs (The VIPs): The big, paid AI models (like GPT-5.5) were the clear winners. They got an average score of 8.67. Think of them as expert security consultants who usually give you the right advice.
  • The Open AIs (The DIY Crew): The free, open-source models (like Llama 4) scored lower, around 6.71. They were "okay," but they made more mistakes, like confusing a medical drug with a hacking term.

3. The "Flip-Flop" Problem: Consistency vs. Quality

Here is the most surprising part of the story. The researchers found that being smart doesn't mean being consistent.

Imagine asking a security guard, "Is this door locked?"

  • Guard A says "Yes" 10 times in a row. (Consistent, but maybe not the smartest guard).
  • Guard B says "Yes" the first time, "No" the second time, and "Maybe" the third time. Even if Guard B is usually right, their flip-flopping is dangerous because you don't know who to trust.

The study found that:

  • The Open AI (Llama 4) was actually the most consistent. It gave the same answer 97% of the time, even though its answers were often lower quality.
  • The Top Commercial AI (GPT 5.5) was the smartest (highest quality), but it occasionally gave contradictory answers. One time it might say, "You should change your password," and the next time it might say, "Your password is fine."

Why This Matters

The paper argues that in the world of security, consistency is just as important as intelligence.

If you are asking an AI how to protect your bank account, you can't afford for it to give you different advice every time you ask. If the AI says "Do X" today and "Don't do X" tomorrow, you might get confused and end up doing nothing, leaving your data exposed.

The Bottom Line

This study is the first to look at real people asking real security questions to AI. It found that:

  1. People use AI for everything from basic learning to trying to hack systems or protect themselves.
  2. The big, paid AI models generally give better advice than the free ones.
  3. However, even the best AI can sometimes contradict itself. To be truly reliable for safety, an AI needs to be both smart and steady.

The researchers conclude that we need to measure both how good an AI's answer is and how consistent it is, because a smart but confused AI is still a risk when it comes to your digital safety.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →