← Latest papers
💻 computer science

Beyond Task Success: An Evidence-Synthesis Framework for Evaluating, Governing, and Orchestrating Agentic AI

This paper addresses the governance-to-action closure gap in Agentic AI by synthesizing evidence from 24 sources to propose a four-layer framework, an ODTA runtime-placement test, and a minimum action-evidence bundle that collectively bridge the disconnect between high-level evaluation/governance and concrete, provable execution compliance.

Original authors: Christopher Koch, Joshua Andreas Wellbrock

Published 2026-04-23
📖 6 min read🧠 Deep dive

Original authors: Christopher Koch, Joshua Andreas Wellbrock

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a very smart, very fast personal assistant named "Agent." This Agent doesn't just answer questions; it has the power to open your bank account, sign contracts, buy supplies, and talk to other people on your behalf.

In the past, we only cared if the Agent finished the job. If you asked it to "buy a laptop," and it came back with a laptop, we said, "Great job!"

But this paper argues that finishing the job isn't enough anymore.

If that Agent bought a $10,000 laptop when your budget was $500, or if it bought it from a scammer because it didn't check the rules, it "succeeded" at the task but failed at being trustworthy. The problem is that we are trying to judge a complex, multi-step process with a simple "pass/fail" grade.

Here is the paper's solution, explained through a simple analogy: The "Smart Chef" in a High-End Restaurant.

The Problem: The "Governance-to-Action" Gap

Imagine a restaurant with a strict Head Chef (the Governance layer). The Head Chef writes a rulebook: "No buying ingredients from unapproved suppliers," and "Never spend more than $50 on a single item."

Then, you have the Agent, who is the sous-chef actually cooking.

  • The Gap: The Head Chef writes the rules, but the sous-chef is running around the kitchen. Sometimes the sous-chef grabs a weird ingredient because it was on sale, thinking, "The Chef won't know."
  • The Issue: We have a rulebook (Governance) and a finished dish (Evaluation), but we are missing the kitchen manager who watches the cooking in real-time to make sure the rules are actually followed while the food is being made.

The paper calls this missing link the "Governance-to-Action Closure Gap." It's the space between "what we said should happen" and "what actually happened in the moment."

The Solution: The Four-Layer Framework

To fix this, the authors propose a four-layer system, like a four-story building where every floor has a specific job:

  1. Layer 1: Evaluation (The Food Critic)

    • Job: Looks at the final dish. Did it taste good? Did the customer get fed?
    • Limitation: The critic only sees the result. They don't know if the chef used poison or stole the ingredients to make it. We need to look at the whole process, not just the plate.
  2. Layer 2: Governance (The Head Chef / Rulebook)

    • Job: Writes the rules. "We only buy organic," "No spending over $50," "Check the supplier's license."
    • Limitation: The rulebook sits on a shelf. It can't stop the sous-chef from grabbing the wrong ingredient in the heat of the moment.
  3. Layer 3: Orchestration (The Kitchen Manager)

    • Job: This is the most important new layer. This is the person standing right next to the stove.
    • Action: When the sous-chef reaches for a $60 steak, the Kitchen Manager says, "Stop! That's over budget. Go get the $40 one instead." Or, "You can't buy from that supplier; I need to call the Head Chef first."
    • Why it matters: This is where the rules actually get enforced in real-time.
  4. Layer 4: Assurance (The Security Camera & Audit)

    • Job: Records everything. If the customer complains later, we can play back the video tape.
    • Action: It proves exactly what happened. "Yes, the Kitchen Manager stopped the expensive steak. Here is the video, the timestamp, and the manager's signature."

The "ODTA" Test: How to Decide Who Does What

The paper introduces a simple test called ODTA to figure out which layer should handle a specific rule. Think of it as a checklist for the Kitchen Manager:

  • O (Observability): Can we see the action before it happens? (Can we see the chef reaching for the expensive steak?)
  • D (Decidability): Is the rule clear? (Is it obvious that $60 is over the $50 limit, or is it a "maybe" situation?)
  • T (Timeliness): Can we stop it fast enough? (If we wait until the steak is cooked, it's too late.)
  • A (Attestability): If we miss it, can we prove what happened later? (Do we have a video recording?)

If the answer is YES to all four: The Kitchen Manager (Orchestration) should stop it immediately.
If the answer is NO: Maybe the rule is too vague for a machine. It should go to a human (Human Judgment) or be checked later by the Security Camera (Assurance).

The "Minimum Action-Evidence Bundle" (MAEB)

Finally, the paper says: "If you are going to let the Agent do something risky (like spending money), you must save a specific 'receipt'."

This isn't just a log saying "I bought a laptop." It's a digital receipt that includes:

  • Who gave the order?
  • What was the exact rule that allowed it?
  • What was the state of the world before the action? (e.g., "We had $40 left in the budget.")
  • Who approved it?
  • A link to the video proof.

Without this "receipt," if something goes wrong, you can't prove who did it or why.

The Real-World Example: The Procurement Agent

The paper uses a "Procurement Agent" (a robot that buys things for a company) to show how this works.

  • Bad Scenario: The robot buys a $10,000 printer. It finishes the task! But it broke the rules.
  • Good Scenario (with the Framework):
    1. Governance says: "Max spend is $5,000."
    2. Orchestration sees the robot trying to buy the $10k printer. It stops the robot and says, "Human approval needed."
    3. A human approves it.
    4. Assurance saves the "receipt" (MAEB) showing the robot tried, the system stopped it, the human approved it, and the money was spent.

The Big Takeaway

We can't just trust AI to "do the right thing" based on its final answer. We need a system where:

  1. Rules are written clearly.
  2. Managers (Orchestration) watch the AI in real-time to enforce those rules.
  3. Recorders (Assurance) save proof of every step.

Trust doesn't come from the AI being smart; it comes from having a system that catches mistakes before they happen and proves they were caught after they happen.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →