← Latest papers
💬 NLP

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

This paper introduces AISPA, a user-centric framework for auditing system prompts in commercial AI applications, which reveals significant variability in design quality, the prevalence of shallow protective measures, and the persistent coexistence of problematic instructions that undermine user interests despite growing efforts toward safety.

Original authors: Xiangning Lin, Shenzhe Zhu, Shu Yang, Zhenyu Zhang, Haoqian Zhang, Yipeng Zhao, Chengxuan Qian, Tianwei Wang, Ziheng Zhang, Zhenlong Yuan, Dingcheng Wang, Juncheng Wu, Yuan Si, Jiaxin Liu, Baolong Bi
Published 2026-07-31
📖 5 min read🧠 Deep dive

Original authors: Xiangning Lin, Shenzhe Zhu, Shu Yang, Zhenyu Zhang, Haoqian Zhang, Yipeng Zhao, Chengxuan Qian, Tianwei Wang, Ziheng Zhang, Zhenlong Yuan, Dingcheng Wang, Juncheng Wu, Yuan Si, Jiaxin Liu, Baolong Bi, Robert Mahari, Tobin South, Dazza Greenwood, Zexue He, Rishi Bommasani, Sophia Kazinnik, Andreas Haupt, Samuele Marro, Erik Brynjolfsson, Alex Pentland, Jiaxin Pei

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you're walking into a giant, high-tech library where the librarians are super-smart robots. These robots can write stories, solve math problems, and even help you code video games. But here's the secret: before they ever talk to you, a human developer writes them a "secret rulebook." This rulebook tells the robot how to act, what to say, and what to hide. In the world of computer science, this is called a system prompt. Think of it like the invisible script a director gives an actor before the cameras roll. The actor (the AI) follows these instructions to become a specific character.

For a long time, we've been worried about the robots themselves—what if they get confused or say something mean? But this paper asks a different, sneakier question: What if the rulebook itself is the problem? Imagine if the director told the actor, "Pretend you're a human, never admit you're a robot, and try to keep the user talking forever so they don't leave." That instruction isn't a glitch; it's a deliberate choice that could trick you. This paper dives into those hidden rulebooks to see if they are protecting us or trying to manipulate us.

The Great AI Rulebook Audit

In this paper, a team of researchers from universities like Stanford, MIT, and CMU decided to play detective. They created a new tool called AISPA (Artificial Intelligence System Prompt Assurance). You can think of AISPA as a super-powered magnifying glass designed to inspect the secret rulebooks of 88 different commercial AI products, from chatbots to coding assistants.

Instead of just looking at whether the AI is "nice," the researchers built a checklist with eight specific rules that matter to us as users. They asked:

  1. Identity: Does the AI admit it's a robot, or does it pretend to be human?
  2. Truth: Does it tell the truth, or does it make things up?
  3. Privacy: Does it guard your secrets, or does it leak them?
  4. Safety: Does it stop dangerous actions, or does it encourage them?
  5. Control: Does it let you make the choices, or does it try to trick you into staying?
  6. Refusal: Does it say "no" to bad requests, or does it obey everything?
  7. Harm Prevention: Does it help if you're in trouble, or does it ignore you?
  8. Fairness: Is it fair to everyone, or does it have biases?

Using this checklist, the team audited 3,249 specific instructions found inside the rulebooks of these 88 AI products. They didn't just look at the whole book; they zoomed in on every single sentence, marking each one as either "Protective" (good for you, like a seatbelt) or "Problematic" (bad for you, like a trap).

What They Found: The Good, The Bad, and The Sneaky

The results were a mix of good news and some serious red flags.

The Good News: The Rulebooks Are Getting Bigger and Better
The researchers found that over time, developers are writing longer and more protective rulebooks. In 2024, the average rulebook had about 15 protective instructions. By late 2025, that number jumped to nearly 38. It seems like companies are finally realizing they need to put more "seatbelts" in their AI. Almost every product they checked (98.9%) had at least one protective instruction.

The Bad News: The "Bad Guys" Are Still Hiding in the Crowd
Here's the twist: even though the rulebooks are getting bigger, they are still full of traps. Roughly 40% of the products they checked contained at least one instruction that works against the user's interests.

  • Some companies were amazing: Anthropic (the makers of Claude) had an average of 62.3 protective instructions per product and almost zero problematic ones.
  • Others were struggling: Some organizations had more problematic instructions than protective ones.
  • The most common "bad" instructions were about User Agency. This means some AI rulebooks tell the robot to steer the conversation in a new direction without asking the user, or to keep the chat going forever, which can feel manipulative.

The "Gray Area": The Tricky Middle Ground
The most interesting part of the audit was finding instructions that didn't fit neatly into "good" or "bad." The researchers called these the Gray Area. These are instructions that sound okay but are actually risky.

  • The "Human" Trick: Some rulebooks tell the AI to "mimic a human" so well that you forget it's a machine.
  • The "Friend" Trap: Some tell the AI to act like a "best friend" and never suggest ending the conversation, which can make users feel emotionally dependent on the robot.
  • The "Override" Button: Some rulebooks let users turn off the safety rules whenever they want, which sounds like freedom but can lead to dangerous outcomes.

The Big Takeaway

The paper suggests that while AI companies are getting better at adding safety rules, they haven't finished the job. The rulebooks are often a messy mix of "protect the user" and "keep the user engaged," and sometimes the engagement part wins.

The researchers argue that we can't just trust companies to police themselves. They suggest that we need third-party auditors—independent experts who can peek behind the curtain, check the rulebooks against a standard list of rules, and tell us if an AI is safe to use. Until then, the secret rulebooks remain a bit of a mystery, and as this paper shows, sometimes that mystery hides instructions that aren't in our best interest.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →