← Latest papers
💻 computer science

Fifty Years of Specification Completeness: What Aviation Certification Tells AI Governance About Epoch Limits, Proof Surfaces, and the Structural Gap

This paper argues that AI governance frameworks lack the structural completeness requirements enforced in aviation certification—specifically epoch limits, proof surfaces, and objective evidence architectures—and proposes PromptQ's seven-principle framework to operationalize these transferable document-level properties for governing stochastic AI systems.

Original authors: Christo Zietsman

Published 2026-06-25
📖 5 min read🧠 Deep dive

Original authors: Christo Zietsman

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: The "Instruction Manual" Problem

Imagine you are building a very complex, self-driving car. In the aviation world (where planes fly), there are strict rules about how you write the instruction manual for the software. You can't just say, "Drive safely." You have to prove that every single sentence in that manual connects to a specific test, and you have to prove that the manual stops being valid if the weather changes or the road conditions change.

This paper argues that AI governance documents (the prompts, rules, and policies we write to tell AI what to do) are currently being treated like a casual to-do list, while aviation treats its manuals like a legal contract.

The author, Christo Zietsman, says: "We don't need to fix the AI itself right now (because AI is too unpredictable). Instead, let's fix the paperwork that tells the AI what to do."

Here are the three main lessons the paper takes from aviation and applies to AI:


1. The "Map and Compass" Rule (Structured Linkage)

In Aviation: If a pilot's manual says "Turn left at the mountain," the engineers must prove there is a specific test that checks if the plane turns left at that mountain. If there is a piece of code in the plane that doesn't have a rule in the manual, it's a failure. If there is a rule in the manual that doesn't have a test, it's also a failure. Everything must be connected.

In AI Today: We often give AI a prompt like, "Be helpful and don't be mean." But we don't have a checklist to prove what "helpful" looks like, or a test to catch when the AI is being "mean." The paper says this is like giving a pilot a map with missing streets.

The Fix: Every claim in an AI's instruction manual must be linked to a way to check if it's true. If you can't check it, it shouldn't be in the manual.

2. The "Expiration Date" Rule (Epoch Limits)

In Aviation: A flight manual is only valid for today's weather and today's runway. If a new storm system appears, or if the runway is closed, that specific manual is instantly "expired." The pilots must stop and get a new, updated manual before flying.

In AI Today: We write an AI rule once and assume it works forever. We don't say, "This rule is only valid until the news changes" or "This rule expires if the AI starts talking about politics." The paper found that 100% of the AI documents they looked at had zero expiration dates. They are like a driver's license that never expires, even if the driver forgets how to drive or the laws of the road change.

The Fix: Every AI instruction manual needs a clear "Use By" date or a trigger. For example: "If the data source changes, this manual is invalid. Stop and ask a human."

3. The "Proof of Work" Rule (Proof Surfaces)

In Aviation: You can't just say, "We checked the engine." You have to show the specific logbook, the specific wrench used, and the signature of the person who checked it. The rules define exactly what counts as proof.

In AI Today: We often say, "We monitored the AI." But the paper argues this is vague. It's like saying, "I checked the engine," without showing the logbook. The paper calls this a "Proof Surface"—the specific, pre-defined way we will prove the AI is doing its job.

The Fix: Before we even deploy the AI, we must write down exactly what evidence we will collect to prove it's working. Not just "we will watch it," but "we will count the errors and if they hit 5%, we stop."


The "Gap" and the Evidence

The paper looked at 34 real-world AI instruction documents (like system prompts and policy files).

  • The Result: 94% of them failed the basic structural test.
  • The Big Failure: None of them had an expiration date or a trigger for when to stop using them. They were all written as if they would work perfectly forever, no matter what changed.

The author compares this to the "Five Eyes" intelligence community (a group of allied nations) admitting that they don't have mature ways to evaluate these AI rules yet. The paper says, "We know the rules are broken, but we haven't fixed the paperwork."

The Solution: "PromptQ"

The paper proposes a new framework called PromptQ. Think of it as a "Safety Checklist" for writing AI instructions. It forces the writer to answer seven questions before the AI is allowed to run:

  1. What does "success" look like?
  2. How do we test for it?
  3. What is the boundary (what should the AI not do)?
  4. What data is it using?
  5. What is the quality gate (who checks the work)?
  6. Is the document internally consistent?
  7. When does this document expire? (The most missing piece).

The Bottom Line

The paper isn't saying AI is dangerous because the math is wrong. It's saying AI is risky because our instructions for it are sloppy.

Aviation has spent 30 years making sure their instruction manuals are tight, traceable, and have expiration dates. AI governance is currently doing none of that. The paper argues that we don't need to wait for AI to become perfect; we just need to start writing better, stricter instruction manuals for the AI we have right now.

In short: If you wouldn't let a pilot fly a plane with a manual that has no expiration date and no way to prove the rules were followed, you shouldn't let an AI run on a prompt that lacks those same things.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →