Fixed-Horizon Self-Normalized Inference for Adaptive Experiments via Martingale AIPW/DML with Logged Propensities
This paper proposes a fixed-horizon inference method for adaptive experiments that leverages logged propensities and predictable nuisance estimation to establish the centered AIPW/DML scores as an exact martingale difference sequence, thereby enabling valid self-normalized confidence intervals for the average treatment effect even when the variance remains random and unpredictable.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are running a massive experiment to see if a new feature on a website (like a red button vs. a blue button) makes people buy more things.
In the old days, you would flip a coin for every visitor: 50% get the red button, 50% get the blue. You wait until the end, count the sales, and calculate the result. Simple.
But in the modern world, companies use "smart" systems (like AI) that learn as they go. If the red button seems to be selling more, the system might start showing it to more people to learn faster. This is called an Adaptive Experiment.
Here is the problem: Because the system keeps changing the rules (the probability of showing the red button) based on what it sees, the math used to calculate the final result gets messy. The "variance" (a measure of how much the results bounce around) becomes unpredictable. It's like trying to measure the speed of a car while the road itself keeps changing its slope.
If you use the standard math formulas (which assume the road is flat and steady), your final confidence interval (your "margin of error") might look correct on average, but it will be wrong for the specific experiment you just ran. You might think you are 95% sure of your answer, but you are actually only 80% sure.
The Paper's Solution: The "Self-Normalizing" Compass
The author, Gabriel Saco, proposes a new way to do the math that doesn't require the road to be flat. He calls it Fixed-Horizon Self-Normalized Inference.
Here is the concept broken down with a simple analogy:
1. The Problem: The Moving Target
Imagine you are trying to hit a bullseye with a bow and arrow.
- Old Way (Standard Math): You assume the wind is steady. You calculate your aim based on a "standard wind speed." If the wind suddenly shifts to a gale, your calculation is wrong, and you miss the target.
- The Adaptive Reality: In these experiments, the "wind" (the assignment probability) changes based on how the arrows are flying. Sometimes it's calm; sometimes it's a storm. You can't predict the wind ahead of time.
2. The Solution: The "Self-Normalizing" Compass
Instead of guessing the wind speed beforehand, the new method says: "Don't guess. Just look at how the arrows actually flew."
The paper suggests using a statistical tool called a Martingale.
- The Metaphor: Think of a hiker walking through a foggy forest. The hiker doesn't know the terrain ahead.
- Standard Math: The hiker tries to use a map that assumes the ground is flat. If the ground turns into a steep hill, the map fails.
- The New Method: The hiker carries a special compass that measures the actual steepness of the ground they have already walked on. With every step, the compass recalibrates itself based on the real terrain. By the time they reach the destination (the end of the experiment), the compass has perfectly adjusted for all the hills and valleys they encountered.
3. The Two Golden Rules (The "Data Contract")
For this "self-calibrating compass" to work, the experiment must follow two strict rules, which the author calls an Auditable Contract:
The "Black Box" Log (Executed Propensities):
The computer system must write down, for every single person, exactly what probability it used to show them the treatment.- Analogy: If the AI decides to show the red button to 80% of people, it must write down "80%" in the log. It cannot just say "I showed it to 80%." It must record the exact number the machine used to make the decision. If the log lies or is missing, the compass breaks.
The "No Peeking" Rule (Predictable Nuisance):
When the system calculates the result, it uses a helper model (a "nuisance" model) to predict outcomes. This helper model must be trained only on data from the past.- Analogy: Imagine a chef tasting a soup. To judge the soup, they must only taste it after it's been served. They cannot peek into the pot and taste the ingredients while they are still cooking, because that would change the flavor. In the experiment, the math model cannot look at the current person's data to make a prediction about them; it must only use data from people who came before.
Why This Matters
- No More "Stabilization": Previous methods tried to force the experiment to behave nicely (e.g., "Don't change the probabilities too much"). This new method says, "Change the probabilities all you want! As long as you log the changes and don't peek, our math will handle it."
- Fixed Horizon: This works perfectly if you decide in advance when to stop the experiment (e.g., "We will stop after 10,000 visitors").
- Warning: You cannot peek at the results and stop early if you see a good result. That breaks the math. You must stick to the pre-agreed finish line.
The Bottom Line
This paper gives experimenters a new, robust toolkit. It allows companies to run smart, adaptive experiments where the rules change on the fly, without breaking the math used to prove if the new feature actually works.
It's like upgrading from a static map (which fails when the terrain changes) to a GPS that updates in real-time based on the actual road conditions, provided you keep a perfect log of where you've been and don't cheat by looking at the future.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.