← Latest papers
🔢 mathematics

Sharp variance estimator and causal bootstrap in stratified randomized experiments

This paper proposes a sharp variance estimator and two randomization-based causal bootstrap methods to improve finite-sample inference for average treatment effects in stratified randomized experiments, addressing the limitations of conservative Neyman-type estimators and normal approximations when sample sizes are small or outcomes are skewed.

Original authors: Haoyang Yu, Ke Zhu, Hanzhong Liu

Published 2026-05-20
📖 6 min read🧠 Deep dive

Original authors: Haoyang Yu, Ke Zhu, Hanzhong Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to figure out if a new medicine (Treatment) works better than a sugar pill (Control). You have a group of people, but they aren't all the same. Some are young, some are old; some live in cities, some in towns. To be fair, you decide to group them into "neighborhoods" (strata) based on these similarities before assigning the medicine. This is a stratified randomized experiment.

Your goal is to measure the "Average Treatment Effect" (ATE)—basically, how much better the medicine makes people feel on average.

The problem is that in the real world, data is messy. Sometimes the sample size is small, or the results are weirdly skewed (like a few people getting extremely good results while most get average ones). Traditional statistical tools used by detectives often get too cautious in these situations. They build a "safety net" (confidence interval) that is so wide it's useless, or they assume the data follows a perfect bell curve when it actually looks like a jagged mountain.

This paper introduces three new tools to help detectives get a sharper, more accurate picture without losing their safety net.

1. The "Sharp" Ruler (Sharp Variance Estimator)

The Old Way: Imagine you are trying to guess the weight of a bag of apples. The old method says, "Let's assume the heaviest possible apple is in the bag, just to be safe." This makes your estimate of the total weight very high, and your "margin of error" huge. You end up saying, "The bag weighs between 10 and 50 pounds." That's safe, but not very helpful.

The New Way: The authors propose a "Sharp Variance Estimator." Instead of guessing the worst-case scenario for every single apple, they use a clever mathematical trick (based on how the apples are ranked) to find the tightest possible safe upper bound.

  • The Metaphor: It's like using a ruler that knows exactly how the apples are stacked. It still keeps the safety net, but it shrinks the "10 to 50 pounds" range down to something like "12 to 18 pounds." It's still safe, but now you actually know what you're dealing with.

2. The "Rank-Preserving" Time Traveler (Causal Bootstrap for Groups)

The Problem: Even with the sharp ruler, if your sample is small or the data is weird, the "bell curve" math (Normal Approximation) can fail. It's like trying to predict the weather using a formula that only works for sunny days, but you're currently in a storm.

The Solution: The authors introduce a "Causal Bootstrap."

  • The Metaphor: Imagine you have a deck of cards representing your experiment. The old method tries to guess the next card by assuming the deck is infinite and perfectly mixed. The new method says, "Let's play the game again and again, but with the exact same cards we have, just shuffled differently."
  • How it works: They use a technique called Rank-Preserving Imputation. Think of it as filling in the missing pieces of a puzzle. If you know Person A got the medicine and felt great, and Person B got the sugar pill and felt okay, this method asks: "What if Person B had taken the medicine? Well, since they are similar to Person A in rank, let's imagine they would have felt just as great relative to their starting point."
  • The Result: By simulating thousands of these "what-if" scenarios using the actual data, they build a map of the uncertainty that is much more accurate than the old bell-curve guess. It's a "second-order refinement," meaning it fixes the small errors the old method misses.

3. The "Constant Effect" Shortcut for Pairs (Causal Bootstrap for Pairs)

The Problem: Sometimes, experiments are so small that people are matched in pairs (like twins). In these "paired experiments," the "Rank-Preserving" method above breaks down. It's like trying to shuffle a deck that only has two cards; there's no room to move them around to create a new scenario.

The Solution: The authors propose a second type of bootstrap specifically for these pairs, based on Constant-Treatment-Effect Imputation.

  • The Metaphor: Imagine you have two twins. One takes a vitamin, one doesn't. The new method assumes a simple rule: "The vitamin adds exactly 5 points of health to anyone who takes it." It doesn't matter who they are; the boost is constant.
  • Why it's clever: Even if the vitamin doesn't actually add exactly 5 points to everyone in real life, this assumption allows the math to work. The authors prove that by using this "constant boost" idea to fill in the missing puzzle pieces, the resulting safety net is still valid and often much tighter than the old methods.

Real-World Tests

The authors tested these tools in two ways:

  1. Simulations: They created fake data with all kinds of weird shapes (skewed, heavy-tailed) and showed their new tools found the truth more often and with tighter ranges than the old tools.
  2. Real Data:
    • Cannabis Cessation Trial: They analyzed a study on quitting marijuana. The data was skewed. Their new method gave a much tighter confidence interval than the standard method, making the results more precise.
    • Mexican Health Insurance: They looked at a study where towns were paired up. The data had outliers (very high costs). Again, their "Constant Effect" bootstrap method provided a more efficient (narrower) confidence interval than the standard approach.

The Bottom Line

This paper doesn't claim to cure diseases or solve poverty. Instead, it provides better mathematical tools for analyzing experiments.

  • If you have groups, use the Sharp Ruler and the Rank-Preserving Time Traveler.
  • If you have pairs, use the Constant Effect Shortcut.

These tools allow researchers to say, "We are 95% sure the effect is between X and Y," with a much tighter, more useful range (X and Y are closer together) than before, without sacrificing the safety of the conclusion. They have even made a free software package (R package CausalBootstrap) so other detectives can use these new tools.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →