← Latest papers
🤖 AI

Principles and Guidelines for Randomized Controlled Trials in AI Evaluation

This paper establishes a foundational framework for standardizing Randomized Controlled Trials in AI evaluation by adapting the Shadish et al. four-validity model and TOP Guidelines into five principles and 33 operational guidelines that prioritize human performance, causal inference, and transparency to address unique challenges in human-AI interaction studies.

Original authors: Christopher Kelly, Angelica Chowdhury, Alexandra Campili, Bimpe Ayoola, Devin Barbour, Thomas Chen Dawson, Ze Shen Chin, Rokas Gipiškis

Published 2026-05-06
📖 5 min read🧠 Deep dive

Original authors: Christopher Kelly, Angelica Chowdhury, Alexandra Campili, Bimpe Ayoola, Devin Barbour, Thomas Chen Dawson, Ze Shen Chin, Rokas Gipiškis

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to figure out if a new, super-smart calculator actually helps people solve math problems better, or if it just makes them feel like they are doing a better job.

For a long time, scientists testing AI have been like chefs tasting their own soup without a recipe. They might say, "This soup tastes great!" but they haven't told you what ingredients they used, how much salt was added, or if they tasted it while wearing a blindfold. Sometimes, the "soup" (the AI) changes flavor halfway through the tasting, or the people tasting it are all professional chefs rather than regular diners.

This paper is a new, strict recipe book for testing AI. It says: "Stop guessing. Let's run a proper, scientific experiment to see if AI actually helps humans."

Here is the simple breakdown of their "recipe," using some everyday analogies:

1. The Core Idea: The "Human Uplift" Test

Most AI tests today are like a car crash test where they just crash the car into a wall and see how much the car breaks. This paper says we need to test the driver, not just the car.

  • The Old Way: "Look, the AI wrote this code!"
  • The New Way: "We gave half the people the AI and half didn't. Did the people with the AI actually write better code, learn faster, or make fewer mistakes than the people without it?"

2. The Five Rules of the Game (The Principles)

The authors borrowed rules from medicine and psychology and added a new one for AI. Think of these as the five pillars holding up a bridge. If one is weak, the bridge (the study) collapses.

  • Rule 1: Are We Measuring the Right Thing? (Construct Validity)

    • Analogy: Imagine you want to test if a new running shoe makes you faster. If you measure "how loud the shoe squeaks," you aren't measuring speed.
    • The Paper's Point: Don't just count how many words an AI helped someone write (that's easy to fake). Measure if the quality of the thinking actually improved.
  • Rule 2: Did the AI Actually Cause the Change? (Internal Validity)

    • Analogy: If you give a group of people a new vitamin and they get healthier, was it the vitamin? Or was it that they also started eating salad and sleeping more?
    • The Paper's Point: You must make sure the only difference between the two groups is the AI. If the "AI group" also got better instructions or a nicer computer, you can't blame the AI for the success. Also, you have to stop the "no-AI" group from secretly sneaking a peek at the AI.
  • Rule 3: Does This Work for Everyone? (External Validity)

    • Analogy: If a medicine works on 20-year-old men in a lab, does it work on 70-year-old women in a hospital?
    • The Paper's Point: Don't test AI only on computer experts or only on university students. If you do, you can't claim it works for everyone. You need to say clearly: "This works for these people, in this situation."
  • Rule 4: Are the Numbers Real? (Statistical Conclusion Validity)

    • Analogy: If you flip a coin 3 times and get heads every time, you might think the coin is magic. But if you flip it 1,000 times, it's probably just luck.
    • The Paper's Point: Don't get excited about tiny improvements that could just be luck. The authors say we need to be super strict with our math (using a stricter "p-value" of 0.005 instead of the usual 0.05) because AI claims are high-stakes. We need to know how much better the AI made things, not just if it made things "better."
  • Rule 5: Show Your Work! (Transparency, Repeatability, and Verification)

    • Analogy: If a magician says, "I made a rabbit appear," but won't let you see the hat or the rabbit, you don't believe him.
    • The Paper's Point: AI changes fast. If you don't write down exactly which version of the AI you used, what prompts you typed, and share your data, nobody can check your work later. The paper creates a 4-level "trust ladder":
      1. Disclosure: "We told you what we did."
      2. Sharing: "We gave you the data and code."
      3. Verification: "An independent person checked that the code actually runs."
      4. Repeatability: "Another team tried it with new people and got the same result."

3. The 33-Step Checklist

The paper doesn't just give rules; it gives a 33-step checklist for researchers.

  • Example: "Before you start, write down exactly what you are testing so you can't change the rules halfway through."
  • Example: "Make sure the people in the 'no AI' group aren't secretly using AI on their own phones."
  • Example: "Check if the AI makes the work faster but worse, or slower but better."

4. Why Do We Need This?

Right now, the world is full of AI studies that are confusing. Some say AI is amazing; others say it's useless. The authors say this is because everyone is playing by different rules.

  • The Problem: If we can't trust the studies, we might deploy dangerous AI, or we might ignore helpful AI.
  • The Solution: This paper is a "standard operating procedure." It's a blueprint for building a bridge of trust. It helps researchers plan their studies, helps reviewers check if a study is good, and helps governments make laws based on real facts, not guesses.

Summary

Think of this paper as the FDA (Food and Drug Administration) for AI experiments. Just as you wouldn't eat a new drug without a rigorous, double-blind clinical trial, we shouldn't trust claims about AI helping humans without these strict, standardized tests. It's about moving from "We think AI is cool" to "We have proof that AI helps these people do this task better."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →