Sample size and power calculations for causal inference of observational studies
This paper establishes a theoretical framework and provides analytical formulas for sample size and power calculations in observational causal inference by decomposing variance into three components and introducing two key parameters—the Bhattacharyya coefficient for covariate overlap and an R-squared-based sensitivity parameter for confounding—supported by a new R package and online calculator.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to figure out if a new medicine actually works. In the perfect world of science, you would run a Randomized Controlled Trial (RCT). This is like flipping a coin for every patient: heads, they get the medicine; tails, they get a sugar pill. Because the coin flip is random, the two groups are identical in every way before the treatment starts. If the medicine group gets better, you know it's the medicine, not something else.
But what if you can't flip coins? Maybe it's unethical, too expensive, or logistically impossible. You have to use Observational Data. This is like looking at a hospital's records where some patients happened to get the medicine and others didn't. The problem? The people who got the medicine might have been sicker, richer, or older to begin with. These differences are called confounders. They muddy the waters, making it hard to tell if the medicine worked or if the patients just had different starting points.
This paper is a guidebook for detectives who have to work with these messy, real-world records. It answers a crucial question: "How many patient records do I need to look at to be sure I'm not fooling myself?"
Here is the breakdown of their solution, using simple analogies:
1. The Problem with Old Maps
Previously, researchers tried to calculate how many records they needed by pretending the messy observational data was a perfect coin-flip experiment. They used standard formulas designed for randomized trials.
- The Flaw: This is like trying to navigate a swamp using a map designed for a highway. It leads to a disaster. You end up thinking you need 100 records, but because the data is so "noisy" (confounded), you actually need 1,000. You run the study, find nothing, and realize you didn't look at enough people.
2. The New Formula: Three Ingredients
The authors developed a new, more accurate formula. They realized that to figure out the "noise" in observational data, you don't need to know every single detail about every patient (which you don't have before the study starts). You only need to understand three specific things:
The Propensity Score Distribution (The "Likelihood" Map):
- What it is: How likely is a person to get the treatment based on their characteristics?
- The Analogy: Imagine a map showing how crowded the "treatment" side of the room is compared to the "control" side. If the groups are completely different (e.g., only young people get the treatment, only old people don't), the map is very uneven. This is called poor overlap.
- The Paper's Trick: They use a special number called the Bhattacharyya coefficient (let's call it the "Overlap Score") to measure how much the two groups look alike. If the score is high, the groups are similar (good). If it's low, they are very different (bad).
The Potential Outcome Distribution (The "Natural Variation"):
- What it is: How much do patient outcomes naturally vary?
- The Analogy: Even without medicine, some patients recover quickly, and some take a long time. This is just natural randomness.
The Correlation (The "Confounding Strength"):
- What it is: How strongly do the patient's characteristics predict both getting the treatment AND the outcome?
- The Analogy: This is the "secret link." If being sick makes you more likely to get the medicine AND makes you less likely to recover, that link creates a huge amount of noise. The authors propose a "sensitivity parameter" (let's call it the "Confounding Coefficient") to measure the strength of this link. They suggest you can estimate this using a standard statistic called R-squared (which measures how well your data predicts the outcome).
3. The "Magic" Result
The paper's biggest breakthrough is showing that you don't need a crystal ball to predict the future data. You only need two extra numbers (the Overlap Score and the Confounding Coefficient) on top of the standard numbers used for randomized trials.
- If the Overlap is good (groups look similar) and Confounding is weak, you need fewer records.
- If the Overlap is bad (groups are totally different) or Confounding is strong, you need many, many more records to get the same level of certainty.
4. The "Safety Net" Approach
The authors built their formula around a specific statistical tool called the Hájek Inverse Probability Weighting estimator.
- The Analogy: Think of this as a "conservative" strategy. It's like packing a parachute that is slightly heavier than necessary. It might not be the lightest, fastest way to fly, but it guarantees you won't fall.
- They chose this method because it doesn't require you to guess the exact shape of the data distribution beforehand. If you guess wrong about the data shape, other methods might fail, but this "heavy parachute" method will still give you a safe, conservative answer (meaning you might calculate you need slightly more people than strictly necessary, but you won't end up with too few).
5. Real-World Testing
The authors tested their formula using:
- Fake Data: They created millions of imaginary patients with different levels of "messiness" (overlap) and "confounding." Their formula consistently predicted the correct number of people needed to find a real effect.
- Real Data: They used a famous dataset about heart catheterization (a medical procedure). They showed that if you used the old "randomized trial" method, you would think you needed about 1,500 patients. But their new method correctly calculated you needed about 3,600. Using the old method would have left the study severely underpowered (likely to miss the truth).
Summary
This paper provides a calculator for observational studies. It tells researchers: "Don't just guess how many records you need. Measure how similar your groups are (Overlap) and how strong the hidden links are (Confounding). Plug those two numbers into our new formula, and you will know exactly how big your study needs to be to avoid wasting time and money."
They even built a free software tool (an R package and an online calculator) so anyone can use this new method without doing the complex math themselves.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.