← Latest papers
📊 statistics

What is estimated in cluster randomized crossover trials with informative sizes? -- A survey of estimands and common estimators

This paper defines and evaluates the consistency of various estimators for multiple treatment effect estimands in 2-period cluster randomized crossover trials with informative sizes, demonstrating that while unweighted and weighted independence estimating equations (and specific fixed or exchangeable models) yield unbiased results, unweighted and weighted nested exchangeable mixed effects models are inconsistent and prone to bias.

Original authors: Kenneth M. Lee, Andrew B. Forbes, Jessica Kasza, Andrew Copas, Brennan C. Kahan, Paul J. Young, Michael O. Harhay, Fan Li

Published 2026-08-12
📖 4 min read☕ Coffee break read

Original authors: Kenneth M. Lee, Andrew B. Forbes, Jessica Kasza, Andrew Copas, Brennan C. Kahan, Paul J. Young, Michael O. Harhay, Fan Li

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery about a new medicine. You can't test it on just one person; you have to test it on whole groups, like different schools, hospitals, or neighborhoods. This is called a "cluster randomized trial." It's like testing a new cafeteria menu by switching the entire lunch line at School A to the new food, while School B keeps the old food, then swapping them the next week. This is a "crossover" design because the groups switch roles.

But here's the tricky part: what if the schools aren't the same size? What if School A has 500 students and School B has only 50? If the new food works great for the big school but poorly for the small one, how do you count the results? Do you care about the average student (the "individual" view), or the average school (the "cluster" view)? In the world of statistics, these are called "estimands"—the specific target you are trying to hit. When the size of the groups changes in a way that is linked to how well the treatment works, statisticians call this "informative sizes." It's like if the biggest schools happened to be the ones where the new food tasted the best just by chance; if you don't account for that, your math might lie to you.

This paper is a guidebook for detectives (researchers) who are solving these mysteries in multi-period crossover trials. The authors, a team of statisticians from universities across the US, UK, Australia, and New Zealand, wanted to figure out which mathematical tools (estimators) give the right answer when the group sizes are "informative." They looked at four main types of tools: simple averages (Independence Estimating Equations), models that treat groups as random samples (Mixed Effects), and models that treat groups as fixed facts (Fixed Effects). They also looked at what happens when you try to "weight" the data to make big and small groups count equally.

The authors ran thousands of computer simulations to test these tools. They found that some popular tools, specifically the "nested exchangeable mixed effects" models, often get the math wrong when group sizes are informative. It's like using a scale that automatically adds extra weight to heavy boxes; if you don't know the boxes are heavy, your final weight is wrong. These models can produce results that look real but actually target a confusing, uninterpretable number that isn't the average student or the average school.

However, the good news is that the "Independence Estimating Equation" (IEE) tools, especially when used with the right weights, are like a reliable, unshakeable scale. They consistently hit the right target, whether you want to know the effect on an individual or a whole group, even when the sizes are messy. The paper also found that "Fixed Effects" models can be tricky; they sometimes accidentally target the wrong thing (like the average time period instead of the average group) if the sizes change in a specific way.

In a real-world test, the authors re-analyzed a study about stomach ulcer medicines for patients on ventilators. They found that the different tools gave slightly different answers, hinting that the hospital sizes might have been "informative." The study concludes that when dealing with these complex, multi-period trials, researchers should be very careful about which tool they pick. If they want to be sure they aren't being misled by the size of the groups, the simple, weighted IEE tools are the safest bet, while the more complex nested models should be used with caution or avoided if the group sizes vary in a way that matters.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →