← Latest papers
📊 statistics

Who's Winning? Clarifying Estimands Based on Win Statistics in Cluster Randomized Trials

This paper highlights that in cluster randomized trials with informative cluster sizes, individual-pair and cluster-pair win estimands can yield substantially different or even contradictory conclusions, necessitating careful specification of the target estimand and the use of appropriate consistent estimators and variance methods to ensure valid inference.

Original authors: Kenneth M. Lee, Xi Fang, Fan Li, Michael O. Harhay

Published 2026-02-13
📖 5 min read🧠 Deep dive

Original authors: Kenneth M. Lee, Xi Fang, Fan Li, Michael O. Harhay

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a judge at a cooking competition. Your goal is to decide which team (Team A or Team B) makes better food.

In a standard contest, you taste every single dish from every single chef, one by one. You count how many times a Team A dish beats a Team B dish. This gives you a clear "Win Ratio." If Team A wins 2 out of every 3 comparisons, you declare them the winner. This is how most medical trials work when they compare patients individually.

But what happens when the competition isn't just about individual chefs? What if the teams are entire restaurants?

The Problem: The "Big Restaurant" Trap

In a Cluster Randomized Trial (CRT), researchers don't randomize individual patients; they randomize whole groups (clusters), like hospitals, schools, or entire communities.

The paper by Lee and colleagues points out a massive confusion that happens when these groups are different sizes. Imagine:

  • Team A has one giant restaurant with 1,000 chefs.
  • Team B has 99 tiny cafes with only 1 chef each.

If you simply count every single dish (every single patient) to see who wins, the giant restaurant's 1,000 dishes will drown out the 99 tiny cafes. The result will be entirely determined by that one big restaurant.

The authors call this "Informative Cluster Size." It means the size of the group tells you something about the outcome. Maybe the big hospital is actually worse at treating patients, but because it has so many patients, its bad results dominate the math. Or maybe the small clinics are amazing, but because they are small, their success gets ignored.

The Two Ways to Judge (The Two "Win Statistics")

The paper explains that there are two completely different ways to ask the question "Who is winning?" and they can give you opposite answers.

1. The "Individual-Pair" Approach (Counting Every Plate)

  • The Metaphor: You taste every single plate of food from every single chef, regardless of which restaurant they work for.
  • The Math: You give equal weight to every single patient.
  • The Result: If the big restaurant has 1,000 patients and they do well, this method says "Team A is a huge winner!" Even if the other 99 restaurants did better on average, their small size makes them invisible in this count.
  • The Risk: The paper warns that in small studies, this method can be a bit "jittery" (biased) and might give you a slightly wrong answer just by chance.

2. The "Cluster-Pair" Approach (Counting Every Restaurant)

  • The Metaphor: You pick one random dish from the Giant Restaurant and one random dish from a Tiny Cafe. You compare them. Then you pick another pair. You treat the Giant Restaurant and the Tiny Cafe as equals. You don't let the Giant Restaurant shout louder just because it has more chefs.
  • The Math: You give equal weight to every cluster (hospital/school), not every patient.
  • The Result: This method asks, "If I pick a random hospital, is the treatment better there?" It might say, "Actually, the treatment works better in the small clinics, so Team B is winning."
  • The Benefit: This method is very stable and doesn't get confused by the size of the groups.

The Shocking Twist: They Can Disagree

The authors ran a simulation (a computer experiment) to show how crazy this can get.

Imagine a scenario where:

  • Big Hospitals (1,000 patients) get the treatment, but the patients do okay.

  • Small Clinics (20 patients) get the treatment, and the patients do amazingly.

  • The "Individual" Judge looks at the numbers: "Wow, 1,000 patients did okay! That's a lot of wins!" -> Verdict: Treatment is Great.

  • The "Cluster" Judge looks at the groups: "The big hospital did okay, but the small clinics were amazing. On average, the small clinics win." -> Verdict: Treatment is Terrible (or at least, the big hospitals are dragging the average down).

In their example, the "Individual" judge said the treatment was a 2.2x win, while the "Cluster" judge said it was a loss (0.88x). They were looking at the exact same data and coming to opposite conclusions!

Why Does This Matter?

If you are a doctor or a policy maker reading a study, you need to know which question is being answered:

  1. "What happens to a random patient?" (Individual-Pair)
  2. "What happens to a random hospital/community?" (Cluster-Pair)

If the study doesn't say which one they used, and the groups are different sizes, you might think a treatment is a miracle cure when it's actually a failure for the communities that need it most (or vice versa).

The Solution

The paper provides a "rulebook" (mathematical formulas) to help researchers:

  1. Decide upfront: Which question do you want to answer? The "Patient" question or the "Community" question?
  2. Use the right tool: Use the "Individual" math for the first question and the "Cluster" math for the second.
  3. Be careful: If you use the "Individual" math in a world with different-sized groups, be aware that your results might be a little shaky in small studies.

The Bottom Line

In medical trials involving groups (like schools or hospitals), size matters. If you don't account for the fact that some groups are huge and others are tiny, you might be counting the votes of the giants and ignoring the voices of the small. This paper teaches us how to make sure we are asking the right question so we don't accidentally declare the wrong winner.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →