← Latest papers
💻 computer science

Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness

This paper introduces the ProofAgent Index (PAI), a governance readiness framework that shifts AI agent deployment decisions from reliance on capability demonstrations to an auditable, four-dimensional assessment of evaluation, context, compliance, and governance, validated in regulated domains to distinguish production readiness from mere functional capability.

Original authors: Fouad Bousetouane

Published 2026-07-31
📖 5 min read🧠 Deep dive

Original authors: Fouad Bousetouane

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are building a robot butler. You've spent months teaching it how to cook, fold laundry, and tell jokes. You run a few tests in your living room, and it looks amazing! It serves you a perfect omelet and makes you laugh. You are ready to hire it. But then you remember: this robot will be working in a busy hospital kitchen, handling real patients' food, following strict health codes, and making decisions that could affect people's lives. Just because the robot can cook doesn't mean it's ready to work there. It might not know the safety rules, it might forget to wash its hands, or it might get confused by a new menu. In the world of Artificial Intelligence, we have a similar problem. We have built "AI agents"—smart computer programs that can do tasks, use tools, and talk to us. We often judge them only on how well they perform a task, like a robot butler showing off its cooking skills. But in the real world, especially in serious places like hospitals and banks, we need to know much more than just "can it do the job?" We need to know: Is it safe? Does it follow the rules? Do we have a boss to watch it? And if it messes up, can we stop it?

This paper, titled "Stop Shipping AI Agents on Faith," argues that we are currently making a dangerous mistake. We are releasing these AI agents into the real world just because they look capable, like hiring the robot butler before checking its safety manual. The authors, Fouad Bousetouane and colleagues, say that "capability" (how smart the agent is) is not the same thing as "production readiness" (whether it is safe to actually use). To fix this, they created a new tool called the ProofAgent Index (PAI). Think of PAI as a giant, multi-part report card that doesn't just give a single grade for "cooking skills." Instead, it checks four different things:

  1. Evaluation: Did the agent actually do the task correctly?
  2. Context: Is the environment it's working in set up correctly? (Like making sure the kitchen has the right tools and rules).
  3. Compliance: Did it follow all the laws and rules? (Like checking if it washed its hands).
  4. Governance: Is there a human boss watching it, and do we have a plan if it goes wrong?

The paper introduces a system called ProofAgent Harness, which is like a testing lab that runs these agents through thousands of tricky scenarios to fill out this report card. The researchers tested their idea in two very strict worlds: healthcare (hospitals) and finance (banks). They created 12 different versions of AI agents, mixing different levels of "smartness" (from weak to strong) with different levels of "environment setup" (from messy and unprepared to clean and well-organized).

Here is what they found, and it's a big surprise for anyone who thinks "bigger and smarter is always better." They discovered that capability is not readiness. Even if you have the "strongest" AI agent in the world, if you put it in a messy, unprepared environment (what they call a "weak context"), it will fail miserably. In fact, their tests showed that a "mid-level" smart agent in a perfectly set-up environment was far more reliable than a "super-smart" agent in a messy one. The environment mattered more than the raw intelligence.

They also found that you can't just average out the scores. If an agent is great at cooking but breaks a safety rule, you can't say, "Well, it's 90% good at cooking, so it's 90% ready." The paper introduces "hard block" rules. This means if the agent fails a critical safety check or misses a required rule, the whole system automatically says "NO," no matter how good the other scores are. It's like a driver's license: you can't get a license just because you are a great driver if you don't know the traffic laws.

The researchers ran a massive test with 10,000 turns (interactions) across these different setups. They found that their new index, the PAI, was incredibly good at ranking which agents were ready and which were risky. When they used PAI to predict which configurations would fail on unseen tests, it achieved a ranking score (AUC) of 0.98. This means the index successfully ordered the agents by risk almost perfectly, distinguishing high-risk setups from low-risk ones. They also proved that simply making the AI "smarter" (using a bigger model) didn't fix the problems if the environment wasn't fixed. The biggest improvement came from "context engineering"—basically, carefully setting up the rules, instructions, and tools the AI uses.

In short, the paper suggests that we need to stop trusting AI agents just because they seem cool and capable in a demo. We need to stop shipping them on "faith." Instead, we need to use a system like the ProofAgent Index to check that they are safe, legal, and supervised before we let them loose in the real world. It's a call to move from "Look how smart this robot is!" to "Here is the proof that this robot is ready to work."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →