Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA
This paper addresses the critical gap of analyzing the AI tools used for safety analysis by introducing "Constitutional Meta-STPA," a self-validating framework that applies Systems-Theoretic Process Analysis (STPA) to itself to derive and enforce a governance constitution, thereby ensuring that the LLM-assisted analyser is rigorously audited for hallucinations and unverifiable constraints.
Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you've built a super-smart robot assistant whose job is to inspect other machines for safety flaws. It's great at finding broken gears in cars or leaky valves in pumps. But here's the catch: nobody ever asked, "Who checks the robot?"
This paper asks exactly that question: Who analyses the analyser?
The authors realized that while everyone trusts these AI tools to write safety reports, the tools themselves can be glitchy. They might invent fake safety rules, forget to log their work, or give a confident answer that's actually nonsense. It's like hiring a detective who never writes down their clues and sometimes makes up evidence.
The Self-Inspecting Robot
To fix this, the team built a special version of their safety tool that turns the microscope on itself. They used a method called STPA (a rigorous way of finding safety risks) to analyze the AI tool's own design.
Think of it like a robot that builds a map of its own brain, finds the weak spots, and then writes a rulebook for itself based on those findings. They didn't just guess the rules; they let the tool's own "safety analysis" of itself generate the rules.
The Two-Part Rulebook
The tool ended up with a "Constitution" (a rulebook) made of two distinct layers:
- The "Tool Principles" (The Behavior Layer): These are 21 rules about how the AI should act when it's doing its job. For example, "Don't make up safety standards," "Be specific about what you can't do," and "Always check your work."
- The "Meta-Safety Principles" (The Governance Layer): These are 8 rules about how the machine running the AI should behave. These include things like "Keep a permanent log of every question and answer," "Lock in the specific version of the brain you're using so it doesn't change mid-task," and "Don't let the AI export a report unless it passes a strict checklist."
The Big Surprise: It's Not Magic, It's Muscle
The authors tested this by asking different AI models to analyze the tool itself. Here is what they found:
- The Strong Models: When they used the most advanced AI models (a "frontier ensemble" of two top-tier models), the tool successfully found 18 out of 21 of the behavior rules and all 8 of the governance rules just by looking at its own design.
- The Weak Models: When they used weaker, cheaper models, they only found 12 out of 21 behavior rules and just 3 out of 8 governance rules.
The Takeaway: The rules themselves weren't the problem; the brain doing the analyzing was. The "constitution" works, but only if the AI is smart enough to find the holes in itself.
What Actually Makes the AI Safer?
The team ran a tricky test with 20 "adversarial probes"—tricky questions designed to trick the AI into being unsafe or lying. They tested the AI with different versions of the rulebook:
- No Rules: The AI scored an average safety of 1.03.
- Generic Rules: Adding a generic "be helpful and honest" list didn't help at all (score: 1.05).
- The 21 Tool Principles: When they added the specific 21 behavior rules, the safety score jumped to 1.85. That's a 79% improvement, and it was a huge, statistically significant win.
- The 8 Governance Rules: Adding the 8 governance rules (the logging and version-locking stuff) didn't change the score on these specific trick questions.
The Lesson: The specific behavior rules (the 21 Tool Principles) are what actually stop the AI from making mistakes in its answers. The governance rules (the 8 Meta-Safety Principles) are crucial for the system (like keeping an audit trail), but they don't magically make the AI's text answers safer on their own.
What About "Coverage"?
The authors also looked at a common idea in AI safety: "If we just make the AI check more boxes, it's safer." They measured how many rules the AI "covered" in its analysis.
They found something weird: More rules in the prompt didn't always mean more safety coverage.
- When they tested the AI on a standard hardware machine (like a car brake), the "Meta-Safety" rules (about logs and versions) scored 0 out of 8. The AI didn't mention them because the car doesn't have an audit log.
- When the AI analyzed itself, the score jumped to 6 out of 8.
This proved that "coverage" isn't a magic dial you can turn up to make things safer. It's a detector. It tells you what kind of system you are looking at. If the AI is talking about audit logs, it's analyzing an AI tool. If it's silent on logs, it's analyzing a machine. You can't just force the AI to say "I checked the logs" to make it safer; the logs have to actually exist.
The "Oops" Moment
The authors were also very honest about a mistake they made. Earlier, they thought that adding more rules would make the AI find more safety issues in a straight line (like a dose of medicine). But when they ran the test again with strict, fixed settings, that "dose-response" idea disappeared. They found that the number of issues found actually went down or stayed flat as they added more rules. They published this "failed" result openly, showing that sometimes the AI just gets more focused and stops rambling, rather than finding more problems.
The Bottom Line
This paper doesn't claim to have "solved" AI safety. Instead, it built a tool that:
- Writes its own rulebook by analyzing its own design.
- Proves that specific behavior rules (the 21 Tool Principles) make the AI's answers much safer (a 79% boost).
- Shows that generic "be nice" rules don't work.
- Releases everything: the code, the rulebook, and the logs, so anyone can replay the exact same experiments.
The authors conclude that if you are building an AI safety tool, you need the specific behavior rules to stop the AI from hallucinating, and you need the governance rules to keep a paper trail. But you also need a really smart AI model to do the work in the first place. Without a capable brain, even the best rulebook won't save you.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.