← Latest papers
📊 statistics

Factorial multistratum designs for building, optimising or evaluating multilevel interventions embedded in learning systems: A taxonomy of factorial cluster-randomised trial designs and their variants

This paper proposes a comprehensive taxonomy for factorial cluster-randomised trial designs, drawing on multistratum concepts from agricultural and industrial research to classify existing clinical examples, illustrate ten design variants using a theoretical obesity intervention, and highlight practical challenges for optimising multilevel healthcare interventions.

Original authors: REA Walwyn, EL Turner, B Copsey, B Goulao, R Foy, AA Montgomery, SG Gilmour, SH Richards

Published 2026-08-03
📖 11 min read🧠 Deep dive

Original authors: REA Walwyn, EL Turner, B Copsey, B Goulao, R Foy, AA Montgomery, SG Gilmour, SH Richards

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to bake the perfect cake, but instead of just mixing flour and sugar, you are trying to fix a whole bakery. You have to decide how to train the bakers, what recipes to put on the menu, and how to arrange the ovens. But here is the tricky part: you can't just test one thing at a time. If you only change the oven temperature, you might miss the fact that the bakers need a different kind of apron to make the cake rise. This is the world of "complex interventions" in healthcare and education. Scientists are trying to figure out how to fix big, messy systems—like hospitals, schools, or entire communities—where many different parts interact. To do this, they use a special tool called a "randomised controlled trial," which is like a scientific experiment where they flip a coin to decide who gets which version of the help. But when you have to test many parts of a system at once, and those parts belong to different groups (like students inside classrooms, inside schools, inside districts), the math gets incredibly complicated. It's like trying to solve a Rubik's cube where every layer is a different color and size, and you have to figure out which twist fixes the whole puzzle without breaking the other pieces.

This paper is a guidebook for untangling that Rubik's cube. The authors, a team of statisticians and researchers, noticed that while scientists have been using these complex, multi-layered experiments for decades in farming and factories, the people running medical and social studies often get confused by the jargon. They call these designs "factorial multistratum designs," which sounds like a mouthful of technical terms, but it basically means testing several things at once across different levels of a system. The paper doesn't invent a new magic pill or claim to have solved every problem in healthcare. Instead, it proposes a new, clearer "taxonomy"—a fancy word for a classification system or a map—to help researchers describe these tricky experiments. By looking at 44 real-world trials and creating a theoretical example about fighting childhood obesity in schools, the authors suggest a standard way to talk about these designs. They argue that if we use this common language, we can better understand how to build, improve, and test interventions that work on multiple levels at once, from the individual patient all the way up to the national policy level.

The Big Picture: Why We Need a Map for Messy Systems

Let's start with the problem. Imagine you are a principal trying to stop kids from gaining too much weight. You could just tell them to eat better. But maybe the problem isn't just the kids; maybe it's the school cafeteria, the gym class, or even the rules the district sets for the whole town. A "multilevel intervention" is when you try to fix all these things at once. You might change the lunch menu (for the school), train the teachers (for the classroom), and run a campaign for parents (for the community).

The trouble is, how do you know which part actually worked? Did the kids lose weight because of the new menu, the teacher training, or the fact that it was a sunny spring? In the old days, scientists might have tested just one thing at a time, like a parallel race where one group gets the menu and another gets nothing. But that's slow and wasteful. It's like trying to find the best car by testing the engine, then the tires, then the paint, one by one, instead of seeing how they work together.

Enter the "factorial design." Think of this as a super-efficient tasting menu. Instead of testing the engine and tires separately, you test four combinations at once:

  1. Old engine, old tires.
  2. New engine, old tires.
  3. Old engine, new tires.
  4. New engine, new tires.

This tells you not only if the new engine is good, but also if the new tires work better with the new engine. This is the "Multiphase Optimisation Strategy" (or MOST), a popular way to build the best possible intervention before trying it on a massive scale.

But here is where it gets messy. In a simple experiment, you flip a coin for every person. In a "cluster-randomised trial" (CRT), you flip the coin for a whole group, like a whole school or a whole hospital. Why? Because if you tell one kid in a school to eat salad and the next kid to eat pizza, the first kid might just grab the second kid's pizza. To avoid this "contamination," you randomise the whole school.

Now, imagine you want to test the menu (for the school) AND the teacher training (for the classroom) AND the parent newsletter (for the district). You can't just flip one coin. You have to flip a coin for the district, a different coin for the school, and another for the classroom. This creates a "multistratum" design. It's like a Russian nesting doll of randomisation. The problem is, the people who study these designs (mostly farmers and factory managers) use a secret language that doctors and teachers don't speak. The textbooks are full of farm examples, not hospital ones. This paper is here to translate that secret language into plain English.

The Paper's Mission: Building a Universal Translator

The authors, led by Rebecca Walwyn and her team, set out to create a common language for these complex trials. They didn't just sit in a lab and dream up theories; they looked at 44 real clinical trials that had already been done. They also built a "theoretical example" based on a real-world scenario: a campaign to reduce childhood obesity in schools.

In their example, they imagined a system with three levels:

  • Districts: The big bosses.
  • Schools: The middle managers.
  • Classrooms: The front line.

They wanted to test four different "ingredients" (treatment factors):

  1. Trainings: For the teachers.
  2. Policies: Rules set by the district.
  3. Menus: What the school serves.
  4. Activities: What happens in the gym.

The paper's main finding is that there are ten distinct ways to arrange these experiments, depending on how you flip your coins and where you apply your ingredients. They created a "taxonomy" (a map) to name and describe these ten designs.

Instead of saying "we did a split-plot thing," which sounds like gardening, they propose a clear three-part description for every trial:

  1. Treatment Structure: What are we testing? (e.g., a 2x2 design means two ingredients, each with two versions: "on" or "off").
  2. Unit Structure: Who is in the experiment? Is it a hierarchy (kids inside classrooms inside schools) or a mix where time periods cross with schools?
  3. Randomisation Design: How did we flip the coins? Did we randomise the whole school, or just the classrooms inside the school?

The Ten Designs: A Tour of the Lab

The authors break down the possibilities into three main families, using their obesity example to show how they work.

Family 1: The Hierarchical Designs (The Nesting Dolls)
These are the most common. Imagine a school system where kids are in classrooms, which are in schools, which are in districts.

  • Design A (The Simple One): You flip one coin for the whole school. The school gets a new menu and new teacher training together. It's a standard "2x2" test, but the whole school is the unit.
  • Design B (The Split-Plot): This is where it gets clever. You flip a coin for the school to decide the menu. But then, inside that same school, you flip another coin for each classroom to decide the teacher training. This is like a "split-plot" design. It's efficient because you don't need to change the whole school's menu for every single classroom, but you can still test the training separately.
  • Design C & D (The Deep Dive): These are even more complex. Design C adds a layer where the outcome is measured on individual students, not just classrooms. Design D adds a third ingredient (district policies) and randomises that at the top level, the school at the middle, and the classroom at the bottom. It's a "split-split-plot" design. It allows you to see exactly which level of the system is doing the heavy lifting.

Family 2: The Cross-Classified Designs (The Time Travelers)
These designs add a twist: Time. Imagine the schools are the same, but the experiment runs for 12 months.

  • Design E: The school gets a treatment (like a new policy) and keeps it for 12 months. You measure the kids every month.
  • Design F: This is a "crossover" design. The school tries Treatment A for 3 months, then Treatment B for 3 months, then A again, then B again. It's like a Latin Square dance where every school tries every combination over time.
  • Design G & H: These mix things up. Maybe the district gets a policy that stays the same for 12 months, but the time periods get a different training every month. Or maybe the school changes its menu every month, but the district policy stays put. This is called a "strip-plot" design. It's great for seeing how a system "learns" over time.

Family 3: The Mixed Designs (The Ultimate Hybrid)
These are the heavyweights. They combine the nesting of Family 1 with the time-crossing of Family 2.

  • Design I & J: Imagine a system where the district policy is fixed for a year, but the school menu changes every month, and the classroom activity changes every week. Or maybe the training changes every month. These designs allow researchers to test four different ingredients across four different levels of the system, all while tracking how the system learns and adapts over time.

What the Paper Says (and Doesn't Say)

The authors are very clear about what they have and haven't done. They suggest that using this new language will make it easier for researchers to plan these trials and for statisticians to analyse them. They found that 44 real-world trials already exist that fit into these categories, proving that these complex designs are not just theoretical—they are being used in primary care, schools, and hospitals right now.

However, the paper does not claim that these designs are always the best choice. In fact, they highlight four big practical challenges:

  1. The System is Messy: Real hospitals and schools aren't perfect boxes. They merge, split, and change size. You can't always get the perfect "balanced" experiment.
  2. Contamination: If you randomise a treatment to a classroom, the kids in the next classroom might find out and copy it. This "spillover" can ruin the experiment, or sometimes, it's actually a good thing (like a vaccine protecting the whole herd).
  3. Carryover: If you test a treatment in January and then switch to a new one in February, the effects of January might still be hanging around in February. This is a risk in time-based designs.
  4. Communication: It's hard to explain to a doctor or a school principal why you need such a complicated design. They might just want a simple answer. The authors note that getting everyone on board requires patience and clear communication.

The Bottom Line

This paper is a toolkit, not a magic wand. It doesn't tell you which intervention will work to stop obesity or cure disease. Instead, it gives researchers a better way to ask the question. By providing a clear map of the ten different ways to set up these multi-level, multi-time experiments, the authors hope to stop the confusion and the "opaque language" that has kept these powerful tools in the shadows.

They suggest that if we use this shared language, we can build better interventions that are tailored to the real world, where problems don't happen in isolation but ripple through families, schools, and communities. The paper suggests that the future of testing complex interventions lies in embracing this complexity, not avoiding it. It's a call to action for the scientific community to stop treating these designs as a mystery and start treating them as the sophisticated, powerful tools they are. The next step, the authors say, is to figure out exactly how big these experiments need to be and how to analyse the data, but for now, we have the map. And with a map, even the most tangled Rubik's cube becomes solvable.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →