Selective Randomization Inference for Adaptive Experiments
This paper proposes a general framework called selective randomization inference that uses conditional post-selection inference and directed acyclic graphs to perform valid statistical testing and construct confidence intervals for adaptive experiments without relying on modeling assumptions or i.i.d. data.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery. In a standard investigation, you decide exactly what questions to ask and how to gather evidence before you start looking at the clues. You stick to that plan, and at the end, you calculate your odds of being right based on that fixed plan.
But in the real world, science often works differently. It's more like a detective who looks at the first few clues, realizes the suspect is hiding in a specific neighborhood, and then decides to focus all their energy on that neighborhood. This is called an adaptive experiment. It's smart because it saves time and resources, but it creates a tricky statistical problem: Selection Bias.
The Problem: The "Unfair" Comparison
The paper explains that if you change your plan based on what you see, you can't just use the old rules to judge your results.
The Analogy of the Dice:
Imagine you are rolling dice to see if a new game is fair.
- Round 1: You roll the dice 10 times. You see that the number "6" keeps coming up.
- The Decision: Because "6" looks promising, you decide, "Okay, from now on, I'm only going to care about the number 6. I'm going to ignore all other numbers."
- Round 2: You roll the dice 10 more times, but this time you only count the "6"s.
If you try to calculate the odds of getting so many "6"s using standard math, you will be wrong. Why? Because you didn't just roll the dice; you chose to focus on "6" because you saw it happen in Round 1. You "cherry-picked" the data. If you compare your results to a standard dice roll (where you didn't pick a number), you are comparing apples to oranges. You are essentially cheating by changing the rules after seeing the first half of the game.
The Solution: "Selective Randomization Inference"
The authors, Tobias Freidling, Qingyuan Zhao, and Zijun Gao, propose a new way to do the math called Selective Randomization Inference.
Think of it as a "What-If" Simulation that respects your detective work.
Instead of asking, "What are the odds of this happening if I had stuck to my original plan?" (which is unfair because you didn't stick to the plan), they ask:
"If I had rolled the dice differently in Round 1, but still ended up deciding to focus on the number 6, how likely would my results be?"
How it works in the paper's language:
- The Setup: They use a "Directed Acyclic Graph" (DAG), which is just a fancy flowchart showing how one thing leads to another (like: See Data Choose Subgroup Assign Treatment).
- The Condition: They only compare your actual experiment to other possible experiments that would have led to the exact same decision to focus on that specific subgroup.
- The Result: This gives you a "Selective P-value." It tells you the odds of your result happening given that you made the specific choice you made.
Why This is Better Than "Data Splitting"
Before this paper, the common advice for this problem was Data Splitting.
- The Old Way: Throw away the first half of your data (the part that made you change your mind) and only use the second half to do the math.
- The Metaphor: It's like a chef tasting a soup, deciding it needs more salt, and then throwing away the first half of the soup to prove the salt works. It's safe, but it's wasteful. You lose a lot of information.
The authors' new method is like Data Carving.
- The New Way: You keep all the data. You just acknowledge that some of the "flavor" (information) was used to make the decision to add salt. You mathematically "carve out" the part used for the decision and only compare the remaining "leftover" information against similar scenarios.
- The Benefit: You get a more powerful test. You use all the data you have, but you don't get fooled by the fact that you made a choice along the way.
The "Hold-Out" Trick
The paper also mentions a practical trick to make the math easier and the results smoother. Sometimes, if you use all the data to make your decision, the math gets jagged and unstable (like a bumpy road).
To fix this, they suggest using a "Hold-Out" group.
- The Metaphor: Imagine you are a teacher grading a class. You look at the first 10 students' tests to decide which topic is hardest. Then, you decide to focus your final exam on that topic. But, to grade the final exam fairly, you only look at the last 10 students' scores, who weren't part of the group you used to make the decision.
- In the paper, they show that if you use a "hold-out" group (data not used for the decision), the math becomes much smoother and more reliable, while still being very powerful.
Summary of Claims
The paper claims that:
- Adaptive experiments are common (in medicine, social science, etc.) but hard to analyze because the rules change based on the data.
- Standard methods fail because they ignore the fact that the researcher made a choice based on the data, leading to false positives (thinking a treatment works when it doesn't).
- Their new method (Selective Randomization Inference) fixes this by only comparing the real experiment to "parallel universes" where the researcher would have made the same choice.
- It requires no assumptions about the data distribution (it doesn't assume the data follows a bell curve or is independent), relying only on the randomness of the treatment assignment.
- It is more powerful than throwing away data (data splitting) because it uses all the available information, just with a more careful mathematical filter.
In short, the paper gives scientists a new, fair rulebook for analyzing experiments where they change their strategy mid-game, ensuring that their conclusions are solid and not just a lucky accident of how they chose to look at the data.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.