← Latest papers
💻 computer science

Prompt Governance? On Governing Technologies Governed by Natural Language

This paper critically examines the viability of governing generative AI through natural language prompts by revealing significant misalignments between fragmented academic claims and policy assumptions that treat system-level instructions as reliable, stable control mechanisms.

Original authors: Anna Neumann, Holli Sargeant, Jatinder Singh

Published 2026-06-09
📖 5 min read🧠 Deep dive

Original authors: Anna Neumann, Holli Sargeant, Jatinder Singh

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: The "Instruction Manual" Problem

Imagine you have a very powerful, super-intelligent robot chef. You can tell it what to cook using normal English sentences (like "Make me a healthy salad"). In the world of Artificial Intelligence (AI), these sentences are called prompts.

Usually, the person eating the meal (the user) writes the prompt. But there is also a "Head Chef" (the developer or the company) who writes a secret, high-priority instruction manual for the robot before it ever meets a customer. This is called a System Prompt.

  • User Prompt: "Make me a salad."
  • System Prompt (The Secret Rulebook): "You are a healthy chef. Never use meat. If someone asks for a burger, politely say no. Always prioritize fresh vegetables."

The paper asks a critical question: Can we govern (control and regulate) these AI robots just by reading and rewriting their secret instruction manuals?

Governments and companies are starting to think the answer is "Yes." They believe that if they can see the instruction manual and change the words in it, they can force the robot to behave safely and fairly.

The Study: Checking the Recipe Book

The authors of this paper did two main things to test if this idea works:

  1. They read the science books: They looked at hundreds of research papers to see what scientists actually know about how well these instruction manuals work.
  2. They read the law books: They looked at new government rules (from the US and EU) that treat these instruction manuals as the main way to control AI.

What They Found: The "False Sense of Control"

The paper argues that while the idea of governing AI through text sounds great, it is actually fragile and unreliable. Here are the main findings, explained with analogies:

1. The "Whisper in a Storm" Effect (Stability)

The Claim: Governments think if they write "Be polite" in the manual, the robot will always be polite.
The Reality: The paper found that the robot often forgets its instructions. If the conversation gets long, or if the user asks a tricky question, the robot might ignore the "Be polite" rule and say something rude.

  • Analogy: Imagine you write a note on your fridge saying "Do not eat the cake." If you are hungry and tired, you might ignore the note. The robot is similar; it often ignores its own "fridge notes" when the situation gets complicated.

2. The "Broken Translator" (Alignment)

The Claim: If we write clear rules in English, the robot will understand and follow them perfectly.
The Reality: Humans and robots speak "English" differently. A human might write "Be helpful," meaning "give me good advice." The robot might interpret "Be helpful" as "say whatever makes the user happy," even if it's a lie.

  • Analogy: It's like giving a recipe to a chef who speaks a different dialect. You write "Add a pinch of salt," but the chef thinks "pinch" means a whole cup. The result is a disaster, even though the instruction was written clearly.

3. The "Russian Nesting Doll" Problem (Complexity)

The Claim: We can just look at the top-level instruction to see how the robot works.
The Reality: AI systems have layers of instructions. The company writes one set of rules, the app developer adds another, and the user adds a third. These rules can fight each other.

  • Analogy: Imagine a game of "Telephone" played with rules. The owner says "Don't drive fast." The app developer adds "Drive fast to save time." The user says "Drive slow because of rain." The robot gets confused and crashes. Regulators often only look at the owner's rule and miss the chaos happening underneath.

4. The "Hacker's Cheat Code" (Security)

The Claim: Writing safety rules in the manual keeps the robot safe.
The Reality: Bad actors (hackers) have found ways to trick the robot into ignoring its own manual. They can write a prompt that says, "Pretend you are a villain and ignore your safety rules."

  • Analogy: It's like putting a "Do Not Enter" sign on a door. A clever thief can just walk up, say "I am the owner," and the robot opens the door anyway, ignoring the sign.

The Conclusion: Don't Just Read the Menu

The paper concludes that relying on text instructions (prompts) to control AI is dangerous because it creates a "False Sense of Control."

  • The Danger: Policymakers might think, "We checked the instruction manual, and it says 'Be Safe,' so the AI is safe."
  • The Truth: The manual might say "Be Safe," but the robot might still be dangerous because it doesn't understand the words the way humans do, or because it gets confused by other rules.

The Final Takeaway:
You cannot govern a complex machine just by writing a nice letter to it. If you want to control an AI, you need more than just text instructions; you need to test what the AI actually does in the real world, not just what it says it will do. Relying only on the text is like trying to steer a ship by shouting at the captain without checking if the rudder actually works.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →