← Latest papers
💻 computer science

Evaluating Prediction-based Interventions with Human Decision Makers In Mind

This paper formalizes various models of human decision-making when aided by predictive systems to demonstrate how cognitive biases and inter-subject dependencies violate standard experimental assumptions, thereby compromising the accuracy of causal effect estimates in automated decision system evaluations.

Original authors: Inioluwa Deborah Raji, Lydia Liu

Published 2026-02-05
📖 5 min read🧠 Deep dive

Original authors: Inioluwa Deborah Raji, Lydia Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to figure out if a new "smart assistant" (like a GPS or a spellchecker) actually helps a human make better decisions. You want to run a test to see if the assistant changes the outcome.

This paper argues that most current tests are flawed because they treat the human decision-maker like a robot that reacts the same way every single time. In reality, humans are messy, emotional, and influenced by how the test is set up.

Here is the breakdown of the paper's main points using everyday analogies:

1. The Problem: The "Robot" Assumption

The Paper's Claim: Most experiments assume that if you show a human a prediction (like a risk score), their reaction is a fixed personality trait. They think, "Judge A always ignores computers, and Judge B always listens."
The Analogy: Imagine testing a new type of coffee. You assume that if you give a cup to Person A, they will always drink it, no matter what. But in reality, Person A might drink it if they are tired, but ignore it if they are already full. Their reaction depends on the context, not just their personality.

The authors say that when we test these "smart assistants" in fields like criminal justice or healthcare, we often forget that the human's reaction changes based on the experiment's design.

2. The Three Ways Humans Get "Tricked" by the Test

The paper proposes that a human's willingness to listen to an algorithm changes based on three specific things the experimenters control. Think of these as three different "moods" the human gets into:

  • Mood 1: The "Familiarity" Effect (Treatment Exposure)

    • What it is: If a human sees the algorithm's advice all the time, they start to trust it and follow it. If they only see it occasionally, they might ignore it or get confused.
    • The Analogy: Imagine a new traffic light. If you see it every day, you eventually stop and go automatically. If it only turns on once a month, you might forget to look at it or think it's broken. The frequency of the light changes how you drive, not just the light itself.
    • The Paper's Point: If an experiment only shows the tool to 50% of cases randomly, the human might not get used to it, leading to a weak result. If they saw it 100% of the time, they might rely on it more, showing a stronger result.
  • Mood 2: The "Overwhelmed" Effect (Capacity Constraint)

    • What it is: If the algorithm screams "DANGER!" (predicts high risk) too often, the human gets tired of the warnings and starts ignoring them, even the real ones.
    • The Analogy: Think of a smoke alarm that goes off every time you toast bread. After a while, you stop listening to it entirely, even when there is a real fire. If the algorithm predicts "high risk" too frequently, the judge stops listening.
    • The Paper's Point: If the experiment sets the algorithm to be very "cautious" (predicting risk often), the human might tune it out, making the tool look useless.
  • Mood 3: The "Cry Wolf" Effect (Low Trust)

    • What it is: If the human sees the algorithm make mistakes, they stop trusting it.
    • The Analogy: If your GPS tells you to turn left, but you know there is a wall there, you will stop trusting the GPS. If the algorithm is wrong often, the human stops following its advice.
    • The Paper's Point: If the experiment uses a model that isn't very accurate, the human will ignore it, making the intervention look ineffective.

3. The Big Mistake: The "Spillover" Problem

The Paper's Claim: In standard science, we assume that what happens to Person A doesn't affect Person B. This is called the "Stable Unit Treatment Value Assumption" (SUTVA).
The Analogy: Imagine testing a new diet pill. You assume that Person A taking the pill doesn't change how Person B reacts to the pill.
The Reality: In this paper, the "pill" is the algorithm. Because the human (the judge) is the same person making decisions for many different cases, what happened to Case #1 changes how they treat Case #2.

  • If they saw the algorithm be right 10 times in a row, they trust it on Case #11.
  • If they saw it be wrong 10 times, they ignore it on Case #11.

This means the decisions are linked. The paper proves that because these decisions are linked, the standard math used to calculate "Did the tool help?" is broken. It often underestimates or overestimates how helpful the tool actually is.

4. The Solution: Change How We Run the Test

The authors ran simulations using real data from a court study (where judges used risk scores). They showed that if you change how you assign the cases:

  • Scenario A: You give the algorithm to 50% of cases for Judge 1, and 50% for Judge 2.
  • Scenario B: You give the algorithm to 100% of cases for Judge 1, and 0% for Judge 2.

Even though the total number of cases is the same, the results change drastically.

  • In the "Familiarity" model, Scenario B (showing it all the time) made the tool look much more effective.
  • In the "Overwhelmed" model, Scenario B made the tool look less effective because the judge got tired of the warnings.

Summary

The paper is a warning to scientists and policymakers: You cannot test a human-AI team the same way you test a machine.

If you want to know if an AI tool actually helps a judge, doctor, or teacher, you have to design your experiment carefully. You have to account for the fact that humans get used to tools, get tired of warnings, and lose trust if the tool makes mistakes. If you ignore these human quirks, your test results will be wrong, and you might think a helpful tool is useless (or vice versa).

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →