A closed-form sample size correction for always-valid inference with optional stopping
This paper introduces a closed-form sample size correction factor that adjusts fixed-sample calculations for always-valid sequential tests, enabling accurate power estimation without simulations and reducing the required sample budget by 8–20% compared to conservative last-point rules.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a case (an A/B test) to see if a new clue (a treatment) is better than the old one. You want to be sure you don't make a mistake, but you also want to stop as soon as you have enough evidence.
The Problem: The "Last-Stop" Trap
In the past, detectives used a rule called the "Last-Stop Rule." They would say, "I will look at the evidence every day, but I will only officially declare a winner on Day 100. I need to make sure that on Day 100 specifically, I have a 95% chance of catching the criminal if they are there."
To be safe, they would calculate how many days they needed to reach that 95% chance. But here's the catch: because they were allowed to peek at the evidence every single day before Day 100, the criminal could have been caught on Day 10, Day 50, or Day 99. The "Last-Stop Rule" ignored all those earlier chances.
Because it ignored the early chances, the rule was being overly cautious. It told the detective, "You need 100 days to be safe," when in reality, they might have been safe with only 85 days. This meant wasting time, money, and resources on a longer investigation than necessary.
The Solution: The "Magic Multiplier"
The author of this paper, Mårten Schultzberg, has found a closed-form correction factor (a fancy math term for a simple "magic number" you can calculate without running complex computer simulations).
Think of this as a magic multiplier (called ). Instead of guessing how long to run the experiment, you take your standard "fixed" calculation and multiply it by this magic number.
- Old Way: "We need 1,000 people to be safe." (This is actually too many).
- New Way: "We need 1,000 people, but multiply that by 0.85. So, we only need 850 people."
This new method accounts for the fact that you are checking the evidence continuously. It realizes that because you are looking all the time, you don't need to wait as long to be sure.
How It Works (The Metaphor)
Imagine the evidence is a ball rolling up a hill.
- The Boundary is a fence at the top of the hill. If the ball crosses the fence, you stop and declare victory.
- The Old Rule assumed the ball only needed to cross the fence at the very end of the track. To be safe, they built a very long track.
- The New Method realizes the ball could cross the fence anywhere along the track.
- The author uses a clever trick: they draw a straight line (a tangent) that just touches the top of the fence at the end point. Because the fence is curved (it gets harder to cross as you go), this straight line sits slightly above the actual fence.
- By calculating the odds of the ball crossing this straight line instead of the curved fence, the math becomes simple and solvable with a calculator. This straight-line approximation is so accurate that it tells you exactly how much shorter the track can be.
The Results: Saving Money and Time
The paper tested this "magic multiplier" against three different types of fences (statistical boundaries) used by companies like Spotify.
- The Savings: By using this new formula, companies can save 8% to 20% of the sample size they would have otherwise collected.
- Analogy: If you were planning to bake a cake for 100 people, this new method tells you, "Actually, you only need ingredients for 85 people, and you'll still be 95% sure the cake is good."
- Accuracy: When they ran thousands of computer simulations, the new method hit the target success rate almost perfectly (within about 3 percentage points).
- Real-World Proof: The author tested this on real data from Spotify's experimentation platform. Across 713 different tests, the method saved a median of 9.5% in sample size. That's a lot of saved money and time.
Important Limits
The paper notes that this "magic number" works best when the data behaves nicely (like a bell curve) and when you have a reasonable amount of data to start with (called "burn-in"). If you are testing something very rare (like a disease that happens to 1 in 10,000 people), the math gets a bit shaky, and the savings might not be as reliable.
Summary
In short, this paper gives experimenters a simple, pre-calculated "discount code" for their sample sizes. It stops them from over-preparing for experiments that allow continuous monitoring, saving them roughly 10–20% of their budget without sacrificing the safety of their conclusions.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.