The AI Evaluability Gap: The Missing Layer for Managing Risk and Sustaining Value
This paper identifies the "AI Evaluability Gap"—the lack of sufficient evidence to support high-confidence governance decisions on risk and value—as a critical flaw in current AI governance, proposing a new framework centered on "Evaluability" that distinguishes between operational and investment certifications based on six properties of evidence to ensure sustainable AI management.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Problem: The "Blind Pilot"
Imagine a company hires a new pilot (an AI system) to fly their plane. The company faces two big questions:
- Is the plane safe enough to fly right now? (Risk Management)
- Is this flight worth the cost of fuel and crew? (Value Management)
Currently, most companies try to answer these questions by looking at the plane's dashboard. But the paper argues that the dashboard is often broken or missing. They don't have enough proof to be confident in their answers.
The authors call this missing proof the "Evaluability Gap." It's not that the AI is necessarily dangerous or useless; it's that the company lacks the evidence to know for sure. Because they can't prove it's safe or valuable, they either fly blindly (taking huge risks) or ground the plane unnecessarily (missing out on value).
The Core Idea: "Evaluability"
The paper introduces a new concept called Evaluability. Think of this not as a feature of the AI itself, but as a feature of the evidence surrounding the AI.
- Old Way: "Is the AI accurate?" (We check the score).
- New Way (Evaluability): "Do we have a trustworthy, up-to-date, and verifiable set of receipts that proves the AI is accurate?"
Evaluability is the ability of a system to generate, keep, and refresh the proof needed to make high-confidence decisions. If you can't prove it, you can't govern it.
The Two Types of Decisions
The paper says we need to stop treating "Safety" and "Money" as the same thing. They require different kinds of proof:
Operational Certification (The "Safety Check"):
- Question: "Can we turn this on?"
- Proof Needed: Structural evidence. Like checking if the brakes work, if the wings are attached, and if the pilot follows the rules.
- Analogy: A car inspection. You need to prove the car won't kill anyone.
Investment Certification (The "Value Check"):
- Question: "Should we keep paying for this?"
- Proof Needed: Causal evidence. Did this specific car actually get us to the destination faster than walking? Did it save us money?
- Analogy: A business expense report. You need to prove the car actually helped the business grow, not just that it exists.
The Trap: A system can be "Operationally Certified" (safe) but fail "Investment Certification" (useless). Currently, companies get stuck because they don't know how to separate these two.
The Six Ingredients of Good Evidence
To close the "Evaluability Gap," the paper says you need six specific ingredients in your evidence. Think of these as the six legs of a sturdy stool. If one is missing, the stool collapses.
- Observability (Can we see it?):
- Analogy: If you are cooking, can you see the ingredients going in and the food coming out? If the AI makes a decision but doesn't log what it saw or what it did, you have no evidence.
- Attributability (Did the AI cause it?):
- Analogy: If sales went up, was it because of the AI, or because it was Christmas? You need to prove the AI caused the result, not just that they happened at the same time.
- Intervenability (Can we stop or change it?):
- Analogy: If the AI starts driving the car into a wall, can you hit the brakes? If you can't turn it off or tweak it to test it, you can't trust it.
- Verifiability (Can someone else check it?):
- Analogy: If you say "I baked a cake," can a friend look at your receipts and the oven logs to prove it? Self-attested evidence (just saying "I did it") isn't enough.
- Calibration (Are we being honest about our confidence?):
- Analogy: If you say, "I am 90% sure this bridge will hold," does it actually hold 90% of the time? Many systems say "90% sure" but are only right 60% of the time. That is dangerous.
- Temporal Validity (Is the proof still fresh?):
- Analogy: A weather report from last week is useless for today. AI changes fast. If the evidence you have is old, it's garbage. You need to keep refreshing the proof.
The "Evidence Lifecycle" (The Cycle of Trust)
The paper argues that evidence isn't a one-time thing. It's like a freshly baked loaf of bread.
- When you bake it (collect evidence), it's fresh and trustworthy.
- Over time, it goes stale (the world changes, users adapt, the AI learns new things).
- Eventually, it becomes moldy (the evidence is no longer valid).
The Solution: You can't just bake the bread once and eat it for a year. You have to keep baking fresh loaves. This means continuous re-certification. You must constantly check if your evidence is still fresh.
Why This Matters for AI Specifically
AI is different from a bridge or a factory machine because AI changes fast.
- The World Moves: Users learn how to trick the AI. Competitors change. The economy shifts.
- The AI Evolves: The AI might learn new things it wasn't trained to do.
- The Result: Evidence that was perfect today might be wrong next week.
Because of this, the paper says we need to stop asking "Is the AI safe?" and start asking "Do we have fresh, trustworthy proof that the AI is safe?"
Summary
The paper concludes that Governance is an evidence problem, not just a technology problem.
- Don't just build better AI models.
- Build better evidence systems around them.
- Make sure you have proof that is observable, attributable, verifiable, calibrated, and fresh.
If you have the proof, you can make confident decisions. If you don't, you are flying blind. The "Evaluability Gap" is the space between "thinking we know" and "having the proof to know."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.