Sample size calculations for multilevel factorial longitudinal cluster randomised trials
This paper presents dedicated sample size methodology for split-plot factorial longitudinal cluster randomised trials with continuous outcomes, enabling the joint assessment of individual-level and cluster-level interventions and their interactions across various longitudinal designs, such as the stepped wedge or cluster randomised crossover.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a chef trying to figure out the perfect recipe for a new dish. You have two main ingredients you want to test: a special herb (let's call this the "Individual Ingredient") and a specific cooking method (let's call this the "Cluster Ingredient").
Usually, scientists run experiments in one of two ways:
- The Individual Test: They give the herb to some people and not others, while everyone cooks the same way.
- The Group Test: They give the cooking method to some entire kitchens (clusters) and not others, while everyone in those kitchens eats the same thing.
But what if you want to test both at the same time? And what if you want to know if the herb tastes better when used with the special cooking method? This is where this paper comes in. It provides a new "mathematical recipe" for figuring out exactly how many people you need to test to get a clear answer.
Here is a breakdown of what the paper does, using simple analogies:
1. The Problem: The "Split-Plot" Kitchen
The paper focuses on a specific type of experiment called a "Split-Plot Factorial Longitudinal Cluster Randomised Trial." That's a mouthful, so let's break it down:
- Split-Plot: Imagine a large garden (the "Cluster"). You decide to water the whole garden with a special fertilizer (the Cluster Intervention). But inside that garden, you have individual plants. You decide to give some of those plants a special spray (the Individual Intervention) and leave others alone. You are testing the fertilizer on the whole garden and the spray on individual plants simultaneously.
- Longitudinal: This isn't just a one-time snapshot. You are watching these gardens over time. Maybe you start with no fertilizer, and then halfway through the season, you switch some gardens to the fertilizer. This is like a "Stepped-Wedge" design, where the intervention rolls out slowly over time.
- The Challenge: When you mix these two things (testing a group change over time and an individual change at the same time), the math gets incredibly messy. Previous math tools could handle the "group over time" part or the "individual" part, but they couldn't easily handle the combination of both, especially when you want to know if the two interventions interact (i.e., does the spray work better with the fertilizer?).
2. The Solution: A New Calculator
The authors (Rhys Bowden and colleagues) have created a new set of closed-form formulas.
Think of previous methods as trying to guess the answer by running thousands of computer simulations (like rolling dice millions of times to see what happens). That takes a long time and doesn't tell you why the answer is what it is.
This paper provides a direct calculator. It gives you a specific equation where you plug in your numbers (how many groups you have, how many people are in each group, how much the people in a group resemble each other), and it instantly tells you:
- How many people you need to recruit.
- How much "power" (confidence) your study will have to detect the effects.
3. The Key Insight: Who Matters More?
One of the most interesting findings in the paper is about where the "noise" comes from.
- The Group Effect: If you want to know if the fertilizer (Cluster Intervention) works, your accuracy depends mostly on how many gardens (clusters) you have.
- The Individual Effect: If you want to know if the spray (Individual Intervention) works, your accuracy depends mostly on how many plants (individuals) you have in total.
The paper shows that even though the fertilizer is applied to the whole garden, the math for the individual spray is surprisingly robust. It scales with the total number of people, not just the number of groups. This is great news for researchers because it means you can get very precise answers about the individual treatment even if you don't have a huge number of groups, as long as you have enough people inside them.
4. The "Interaction" Twist
The paper also tackles the tricky question of interaction.
- Does the spray work better when the fertilizer is present?
The authors show that in these specific "split-plot" designs, you can actually estimate this interaction very precisely. In fact, as you add more people to each time period, your ability to detect this "interaction" becomes even sharper than your ability to detect the main group effect. It's like having a magnifying glass that gets stronger the more people you add to the study.
5. Real-World Example: The SharES Trial
To prove their math works, the authors applied it to a real study called the SharES trial.
- The Setup: This trial looked at breast cancer patients.
- Individual Level: Patients got a decision-making tool (like a pamphlet or app).
- Group Level: Clinicians got a dashboard to help them review patient needs.
- The Result: The authors used their new formulas to calculate exactly how many patients needed to be in each clinic at each time period to see if the tools worked. They showed that if you ignore the interaction between the two tools, you need fewer people. But if you do want to see if the tools work better together, you need to recruit more people.
Summary
In short, this paper is a instruction manual for building better experiments. It tells researchers exactly how to size their studies when they are testing two different things at once—one that affects whole groups over time, and one that affects individuals. It replaces slow, guesswork-heavy simulations with fast, clear math, ensuring that studies are neither too small (wasting time) nor too big (wasting money).
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.