← Latest papers
💻 computer science

HelpBench: Assessing the Ability of LLMs to Provide Privacy, Safety, and Security Advice

This paper introduces HelpBench, a benchmark comprising 450 authentic questions and evaluation rubrics that reveals while state-of-the-art LLMs generally provide high-quality advice on digital privacy, safety, and security, a significant portion of their responses still contain inaccurate or potentially harmful information.

Original authors: Sarah Meiklejohn, Sunny Consolvo, Patrick Gage Kelley, Tara Matthews, Sai Teja Peddinti, Renee Shelby, Lenin Simicich, Kurt Thomas

Published 2026-06-24
📖 5 min read🧠 Deep dive

Original authors: Sarah Meiklejohn, Sunny Consolvo, Patrick Gage Kelley, Tara Matthews, Sai Teja Peddinti, Renee Shelby, Lenin Simicich, Kurt Thomas

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, well-read librarian who knows a little bit about everything. You go to this librarian and ask, "My computer is acting weird, and I think I'm being watched. What should I do?" or "I got an email saying I won a prize, but it looks fake. Is it safe?"

This paper, called HelpBench, is like a giant report card for these "librarians" (which are actually Large Language Models, or LLMs) to see how well they answer questions about digital privacy, safety, and security.

Here is the story of what the researchers did and what they found, explained simply:

1. The Problem: A Dangerous Library

People are starting to ask these AI librarians for help with serious problems, like hacked accounts, scams, or abusive stalkers. But there's a catch: if the librarian gives you the wrong advice, you could lose your money, your identity, or even your physical safety.

The researchers wanted to know: Can these AI librarians be trusted with these high-stakes questions?

2. Building the Test: The "HelpBench"

To test the librarians, the team created a special exam called HelpBench.

  • The Questions: They didn't just make up fake questions. They went into a giant digital town square (Reddit) and found 450 real questions that regular people had asked when they were in trouble. They cleaned up the questions to protect people's privacy but kept the real-life stress and confusion intact.
  • The Topics: The questions covered nine scary areas, like:
    • Getting back into a locked account.
    • Figuring out if an email is a scam.
    • Dealing with harassment or stalking.
    • Protecting your data from being stolen.
  • The Answer Key: For every single question, a team of human experts (people with over 10 years of experience in digital safety) wrote a detailed "answer key." This key didn't just say "Right" or "Wrong." It checked:
    • Facts: Did they give the correct technical advice? (e.g., "Don't click that link!")
    • Tone: Was the advice calm and clear, or did it panic the user?
    • Safety: Did they accidentally suggest something dangerous?

3. The Exam: Testing 18 Librarians

The researchers took 18 of the smartest AI models available today and asked them all 450 questions. That's 40,500 answers in total! They used a computer program (an "auto-rater") to grade the answers against the human experts' keys.

4. The Results: Good News, Bad News

The results were a mix of impressive skill and dangerous mistakes.

  • The Good News: On average, the AI librarians did a pretty good job. They scored about 82% overall. For simple questions, like "Is this email a scam?", they were often very accurate.
  • The Bad News: The average score hides a scary reality. One out of every ten answers was a failure. These weren't just "oops, I forgot a comma" mistakes. These were answers that scored below 65%, meaning they gave inaccurate or even harmful advice.

5. What Went Wrong? (The "Long Tail" of Failure)

The paper highlights some specific ways the AI failed, which are like a librarian giving you the wrong map in a dangerous neighborhood:

  • Missing the "At-Risk" Context: If a user asked about removing spyware from a computer used by an abusive partner, the AI might just give technical steps. It failed to realize that suddenly removing the spyware could trigger the abuser to become violent. The AI missed the human danger.
  • False Reassurance: In one case, a user asked if adding a file to a "vault" deleted the original copy. The AI said yes, but it didn't. This gave the user a false sense of security, leaving their data vulnerable.
  • Bad Advice for the Poor: Sometimes the AI suggested solutions that were too expensive or impossible, like "buy a brand new computer" or "open a bank account at a different bank" just to fix a simple app issue.
  • Encouraging Rule-Breaking: When users asked how to get around a ban on a social media site, some AIs said, "Well, that's between you and the app," effectively encouraging them to break the rules and get banned again.
  • Reading Minds: The AI sometimes tried to guess why someone blocked them, inventing dramatic stories about the other person's motives, which can make the user feel more anxious.

6. The Conclusion

The paper concludes that while these AI tools are getting better, they are not yet ready to be the sole source of truth for digital safety.

Think of it like a student pilot. They can fly the plane on a clear day (average questions), but if they hit a storm (complex safety issues) or a tricky landing (high-risk situations), they might crash. Because the stakes are so high (losing your money, your privacy, or your safety), the researchers say we need to keep testing and improving these models until they are consistently perfect, not just "mostly" good.

In short: AI is a helpful assistant, but when it comes to your digital safety, you shouldn't trust it with your life just yet. It still makes dangerous mistakes too often.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →