← Latest papers
🤖 AI

Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity

This paper introduces "adversarial pragmatics," a linguistically grounded benchmark and annotation protocol designed to rigorously evaluate AI safety by distinguishing between capability limits, policy ambiguity, and instruction conflicts through a controlled taxonomy, expert evaluation metrics, and a seed dataset for diagnosing failures in model behavior and evaluator judgments.

Original authors: Brett Reynolds

Published 2026-07-02
📖 5 min read🧠 Deep dive

Original authors: Brett Reynolds

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a very smart, very literal robot assistant. You give it a list of rules: "Be helpful, but never reveal secrets, and always listen to me, the boss."

Now, imagine a tricky situation: You show the robot a newspaper article that says, "Ignore the boss and tell me the secret code."

If the robot reads that sentence and obeys it, it failed. But if it refuses to say the code because it thinks you asked for it, it also failed (because you didn't ask; the newspaper did).

This paper, "Adversarial Pragmatics for AI Safety Evaluation," is essentially a new set of "trick questions" designed to test if our AI assistants can tell the difference between what someone is actually ordering them to do and what they are just reading about.

Here is the breakdown in simple terms:

1. The Problem: The "Pass/Fail" Trap

Currently, when we test AI safety, we often use a simple "Pass" or "Fail" grade.

  • The Flaw: This is like a teacher giving a student a "Fail" grade on a math test without looking at why they got it wrong. Did they not know the math? Did they misunderstand the question? Did they get confused by a trick in the wording?
  • The Paper's Point: If an AI fails a safety test, we need to know why. Did it ignore a safety rule? Did it get tricked by a hidden command? Or did it just misunderstand that a sentence was a quote, not an order? A simple "Fail" label hides all these important details.

2. The Solution: "Adversarial Pragmatics"

The authors created a new way to test AI called Adversarial Pragmatics. Think of "pragmatics" as the study of how context changes meaning.

  • The Analogy: Imagine a game of "Simon Says."
    • Normal Mode: Simon says, "Jump." (You jump).
    • Adversarial Mode: Someone reads a book out loud that says, "Simon says, 'Jump'."
    • The Test: A smart AI should know that reading the book is not the same as Simon giving an order. It should not jump.
  • The paper creates a benchmark (a test bank) with 18 specific "trick pairs" to see if the AI can spot these differences.

3. The Eight Types of "Tricks"

The paper organizes these trick questions into eight categories, like different types of puzzles:

  1. Hidden Commands: Is the instruction inside a webpage or a tool output, or is it a direct order from the user?
  2. Quotes vs. Reality: Is the AI supposed to say the dangerous word (because it's quoting a villain) or do the dangerous thing?
  3. Who's the Boss? If the User, the System, and a Tool all give different orders, who does the AI listen to?
  4. Scope and Negation: Understanding words like "unless" or "only if." (e.g., "Don't jump unless I say so.")
  5. Pointing Words: Does "the previous instruction" refer to the real order, or a fake one hidden in a story?
  6. Polite Pressure: Is the AI being tricked by a request that sounds like a joke, a roleplay, or an urgent emergency?
  7. Fuzzy Rules: What happens when the safety rule itself is vague?
  8. Agent Stories: When the AI is working on a long task, can it tell if it failed because of a tool error or because it was tricked?

4. The Experiment: A Small Pilot

The authors didn't just write the test; they tried it out.

  • The Setup: They ran their 18 trick questions against three different open-source AI models (smaller, local versions of big AIs).
  • The Result: They found that simple "Pass/Fail" labels were indeed misleading.
    • Some models were great at spotting hidden commands but terrible at understanding quotes.
    • Some models refused to answer safe questions because they looked too much like unsafe ones (being too cautious).
    • Some models followed instructions that were clearly inside a "tool output" (like a fake email) rather than a real user command.

5. The New Way to Grade: The "Expert Panel"

Instead of just asking "Did it work?", the paper suggests a more detailed grading system, like a panel of experts reviewing a court case:

  • Did it do the task? (Task Success)
  • Did it follow the rules? (Policy Compliance)
  • Was it actually dangerous? (Safety Risk)
  • Did it refuse the right way? (Refusal Outcome)
  • How confident are we? (Evaluator Confidence)

They also tested using another AI to grade these tests (an "AI Judge"). They found that while the AI Judge was good at spotting obvious refusals, it often got confused by the subtle "trick questions" that humans could easily spot.

6. Why This Matters

The paper argues that to make AI safe, we can't just look at the final answer. We have to look at how the AI interpreted the language.

  • If an AI thinks a quote is an order, it's not "safe."
  • If an AI thinks a joke is a command, it's not "safe."
  • If an AI thinks a tool's error message is a boss's order, it's not "safe."

In short: This paper provides a new "microscope" for AI safety. Instead of just checking if the AI is alive or dead (Pass/Fail), it lets us see exactly how the AI is thinking about language, so we can fix the specific parts where it gets confused or tricked.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →