← Latest papers
🤖 AI

Towards AI epidemiology: a measurement standardisation framework for prospective risk detection

This concept paper proposes a measurement standardisation framework to enable "AI epidemiology" for prospective risk detection in deployed AI systems by defining a grammar and statistical protocol to validate the reliability of expert-AI interaction assessments and establish a basis for future governance and epidemiological studies.

Original authors: Kit Tempest-Walters

Published 2026-06-05
📖 5 min read🧠 Deep dive

Original authors: Kit Tempest-Walters

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are the manager of a busy hospital, a bank, or a law firm. You've hired a team of incredibly smart, fast, but invisible assistants (AI models) to help your experts make decisions. The problem is, these assistants are "black boxes." You can see what they say, but you have no idea how they think inside. Sometimes they give great advice; sometimes they give dangerous or rule-breaking advice, and you don't know why until it's too late.

This paper proposes a new way to manage these invisible assistants without needing to open up their "brains." It suggests a system called AI Epidemiology.

Here is the breakdown of the paper's ideas using simple analogies:

1. The Problem: We Can't See Inside the Machine

Currently, scientists try to understand AI by looking at its internal code (like trying to understand a car engine by taking it apart). The paper calls this "correspondence-based interpretability." But as AI gets bigger, this becomes impossible. The engine is too complex, and the parts are too tangled.

The Paper's Solution: Instead of trying to see how the engine works, let's just watch what the car does and how the driver reacts to it. We don't need to know the mechanics to know if the car is driving safely.

2. The Core Idea: The "Traffic Cop" System

The authors propose a framework that acts like a Traffic Cop standing next to every interaction between a human expert and the AI.

Every time an expert asks the AI a question and gets an answer, the system automatically records three things in a standardized "report card":

  • The Mission: What was the question? (e.g., "Should I approve this loan?")
  • The Conclusion: What did the AI say? (e.g., "Reject the loan.")
  • The Justification: Why did the AI say that? (e.g., "The credit score is too low.")

The system then gives this report card two scores:

  1. Policy Alignment: Does this answer follow the company's rules?
  2. Evidential Alignment: Is the AI's reasoning based on real facts, or is it making things up?

3. The "Judge" Who is Also an AI (The Circular Problem)

You might ask: "Who grades the report card?"
The paper admits a funny problem: They use another AI (a Large Language Model) to grade the first AI's answers. It's like hiring a second invisible assistant to check the work of the first one.

How they fix this: They don't just let the second AI guess. They give it a strict Rubric (a checklist), force it to look at the official rulebooks (Reference Documents), and make it think step-by-step before giving a score. They also run tests to make sure the "Judge AI" isn't just being a "yes-man" (agreeing with everything) or liking long answers just because they are long.

4. The "Epidemiology" Analogy

This is the most creative part of the paper. The authors compare their system to Public Health Epidemiology.

  • The Old Way (Mechanism): In the past, doctors didn't know why smoking caused cancer. They didn't know the molecular biology. But they noticed a pattern: Everyone who smoked got sick. They didn't need to understand the biology to tell people to stop smoking.
  • The New Way (Correlation): Similarly, this framework doesn't need to know how the AI makes a mistake. It just needs to notice patterns.
    • Example: "Every time the AI talks about 'post-bankruptcy loans,' it gets a low score on facts."
    • Result: The institution knows, "Hey, our AI is bad at this specific topic," and they can fix it or have humans double-check those specific cases.

5. What Happens in Real Life?

The paper describes a workflow for the human expert:

  1. The expert asks the AI a question.
  2. The AI answers.
  3. Instantly, the system shows the expert a "Traffic Light" signal:
    • Green: High scores on rules and facts. Go ahead.
    • Yellow/Red: Low scores. The system flags it.
  4. If the light is red, the expert is prompted to hit an "Override" button. They can then correct the AI's answer.
  5. The system quietly saves all these "Traffic Light" data points.

6. The Big Picture: Three Stages of Proof

The paper doesn't claim to have solved everything yet. It outlines a three-stage research plan to prove this works:

  • Stage 1 (Reliability): Prove that the "Judge AI" can consistently grade the answers correctly, just like a human would.
  • Stage 2 (Governance): Prove that these grades actually help experts and managers spot trouble before it happens.
  • Stage 3 (Outcome): Prove that when the AI gets a "Red Light," it actually leads to bad real-world results (like a bad loan or a medical error), confirming that the system is detecting real risks.

Summary

This paper proposes a standardized reporting system for AI. Instead of trying to understand the AI's secret thoughts, it treats AI interactions like disease symptoms. By collecting thousands of these "symptoms" (standardized reports) and grading them, institutions can spot dangerous patterns and fix them before they cause harm, all without needing to see inside the AI's black box.

Important Note: The paper explicitly states that this is a proposal and a framework. It does not claim to have already solved AI safety or to have results from real-world hospitals or banks yet. It is setting the rules for how we will test this in the future.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →