← Latest papers
📊 statistics

Hypothesizing an effect size by considering individual variation

This paper proposes a method for formulating realistic hypotheses about average treatment effects in experimental and observational studies by first conceptualizing the underlying distribution of individual effects, illustrated through examples from medicine, economics, and psychology.

Original authors: Andrew Gelman, Amy Krefman, Lauren Kennedy, Jessica Hullman

Published 2026-04-10
📖 6 min read🧠 Deep dive

Original authors: Andrew Gelman, Amy Krefman, Lauren Kennedy, Jessica Hullman

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Stop Guessing the "Average," Start Imagining the "Crowd"

Imagine you are a chef planning a new restaurant. You want to know how much money you'll make.

The Old Way (The Paper's Critique):
Most researchers act like they are guessing a single number. They say, "I think my new dish will make $100 profit per day." They base this on a guess, a tiny test run, or a story they heard from a friend.

  • The Problem: If they guess $100, but the real average is only $10, they will buy too many ingredients, hire too many staff, and go bankrupt. In science, this leads to studies that are too small to find real answers, or studies that find "fake" big answers just by luck.

The New Way (The Paper's Solution):
Instead of guessing one number, the authors say: Stop thinking about the "average person" and start thinking about the "whole crowd."

They propose a simple mental shift: Don't ask, "What is the effect?" Ask, "Who does this help, who does it hurt, and who doesn't care at all?"


The "Recipe" for a Better Guess

The authors suggest a four-step recipe to figure out what a realistic effect size looks like. Think of it like planning a party:

1. Imagine the "Super Fans" (The Plausible Effect)

First, imagine the people for whom your treatment works perfectly.

  • Analogy: If you are testing a new study app, imagine the student who is already motivated, has a quiet room, and loves technology. For them, the app might boost their grades by 20%.
  • Action: Start with this big, optimistic number. Let's call it X.

2. Imagine the "Real World" (The Range)

Now, zoom out. Not everyone is a "Super Fan."

  • Analogy: Some students are distracted, some hate technology, and some are just having a bad week. For them, the app might do nothing, or even make them stressed (a negative effect).
  • Action: Realize that the effect isn't just "20%." It's a mix. Maybe for some it's 20%, for others it's 0%, and for a few it's -5%.

3. The "Null" Group (The People Who Don't Care)

This is the most important step. A huge chunk of people simply won't be affected at all.

  • Analogy: If you are testing a new flavor of ice cream, 50% of people might not eat ice cream at all. If you are testing a medicine, some people might not take the pill, or their bodies might not absorb it.
  • Action: Assume a big chunk of your crowd (maybe 50% or even 90%) gets zero effect.

4. Do the Math (The New Average)

Now, combine these groups.

  • If your "Super Fan" boost is X, but only 50% of people are Super Fans, and the other 50% get 0, your real average isn't X. It's X divided by 2.
  • If you are in marketing (where people ignore ads), maybe only 10% care. Then your average is X divided by 10.

The Result: You realize that your "big" guess of $100 profit was actually more like $10. Now you can plan your study (or your restaurant) with a realistic budget.


Why This Matters: Three Real-World Examples

The paper uses three stories to show why this matters:

1. The Medical Miracle (Binary Outcomes)

  • Scenario: A doctor thinks a new drug saves lives. He guesses it saves 25% more people.
  • The "Crowd" View: The authors ask: Who actually gets saved?
    • Some people would have lived anyway (0% benefit).
    • Some people are too sick to be saved (0% benefit).
    • Only a small group is "on the fence" and gets saved (100% benefit).
  • The Lesson: If the drug saves 25% of the total group, it might actually be saving 100% of a very small group. If the doctor designs a study for a "general" population, he might need 16 times more patients than he thought to see that small average effect.

2. The Political Opinion (The "Middle" Problem)

  • Scenario: Researchers want to know if knowing a gay person changes your opinion on marriage.
  • The "Crowd" View:
    • 1/3 of people are already strongly against it (won't change).
    • 1/3 are already strongly for it (won't change).
    • Only 1/3 are "on the fence."
    • Of that 1/3, maybe half don't change, and only a few shift their opinion.
  • The Lesson: The "average" change across the whole country is tiny. If researchers expect a huge shift, they will be disappointed. They need to design a study that expects a tiny signal, not a loud one.

3. The "Too Good to Be True" Study

  • Scenario: A study claims a childhood program increases adult earnings by 42%.
  • The "Crowd" View: The authors say, "That's impossible as an average."
    • For the effect to be 42% on average, it would mean everyone got a massive raise, or a few people got a 1000% raise while everyone else got nothing.
    • In reality, most people get a tiny boost, and some get nothing.
  • The Lesson: When you see a study with a "huge" average effect, it's often a red flag. It usually means the study was too small, got lucky, or only looked at the "Super Fans" (the people who would have succeeded anyway).

The "Aha!" Moment for Scientists

The paper argues that the scientific community has been suffering from "Bayesian Cringe." This is a fancy way of saying: "We are so afraid of being wrong that we refuse to make any assumptions, so we end up assuming effects are huge and easy to find."

The Fix:
Instead of trying to find a single "Truth" (e.g., "This drug works"), we should accept that effects vary.

  • Some people benefit.
  • Some don't.
  • Some are harmed.

By mapping out this variation before you start your experiment, you stop overestimating your power. You stop designing studies that are doomed to fail. You stop getting fooled by "lucky" results that look big but aren't real.

Summary in One Sentence

Don't guess the size of the wave; instead, imagine the whole ocean, figure out how many people are actually swimming, and then calculate how big the wave really is for the people who are in the water.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →