← Latest papers
📈 economics

Randomization Tests in Randomized Saturation Designs

This paper develops finite-sample valid and asymptotically justified randomization tests for various null hypotheses regarding spillover effects in randomized saturation designs, including individual-level, average, and monotonicity assumptions, and demonstrates their practical utility through simulations and a real-world application.

Original authors: Jizhou Liu, Azeem M. Shaikh, Liang Zhong

Published 2026-07-07
📖 5 min read🧠 Deep dive

Original authors: Jizhou Liu, Azeem M. Shaikh, Liang Zhong

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to figure out if a new teaching method works. But there's a twist: you can't just test it on individual students in isolation. If Student A learns the method, they might help Student B learn it too, even if Student B wasn't officially assigned the method. This is called a spillover effect.

To study this, researchers use a Randomized Saturation Design. Think of this like a school district experiment:

  1. The Clusters: Instead of testing students one by one, you test whole classrooms (clusters).
  2. The Saturation: You assign different classrooms to have different "densities" of the new method.
    • Classroom A gets 0% of students using the method.
    • Classroom B gets 33% of students using it.
    • Classroom C gets 66%.
    • Classroom D gets 100%.
  3. The Goal: You want to see if the students who didn't use the method (the untreated ones) still learned more just because their classmates did.

The Problem:
Usually, statisticians use big, complex math formulas to guess if these results are real or just luck. But these experiments often have very few classrooms (maybe only 10 or 20). When you have such a small group, those big formulas often break down and give wrong answers. Also, the classrooms are often different sizes, making the math even messier.

The Solution: The "What If" Game (Randomization Tests)
The authors of this paper propose a smarter way to play the "What If" game. Instead of guessing with formulas, they simulate thousands of alternative realities based only on the rules of the experiment.

Here is how they do it, broken down by the three main tools they invented:

1. The "Focal Unit" Filter (For Specific Comparisons)

Imagine you want to compare the 33% classroom against the 0% classroom.

  • The Trick: You don't look at every student in the room. You pick a few specific "focal" students who didn't use the method in both classrooms.
  • The Game: You pretend to swap the labels of the classrooms. You imagine, "What if Classroom A was actually the 0% room and Classroom B was the 33% room?"
  • The Result: Because you only looked at students who didn't use the method, and you know the rules of the game, you can calculate exactly how likely it is that the results you saw happened by pure chance. This gives a perfectly accurate answer for small groups, no matter how weird the data looks.

2. The "Studentized" Score (For Average Effects)

Sometimes, you don't care if every single student is affected; you just want to know if the average student in the 33% room did better than the average in the 0% room.

  • The Challenge: The "perfect" math trick above doesn't work for averages because the "What If" game gets messy when you only care about the average.
  • The Fix: The authors created a special "scorecard" (a studentized statistic) that adjusts for how much the data naturally varies.
  • The Result: While this isn't a "perfect" answer for a tiny group, it becomes extremely accurate as you add more classrooms. It bridges the gap between the small-group "perfect" math and the big-group "approximate" math.

3. The "Monotone" Check (For Multiple Levels)

What if you have four levels (0%, 33%, 66%, 100%) and you want to know: "Does the effect get stronger as we add more students using the method?"

  • The Old Way: You would have to run three separate tests (0 vs 33, 33 vs 66, 66 vs 100) and hope they all line up. This is like checking if a ladder is straight by measuring each rung separately; it's prone to errors.
  • The New Way: The authors built a "Global Monotone Test." It looks at the entire ladder at once. It asks, "Is there any way to rearrange these classrooms that would make the results look like a straight, upward-sloping line?"
  • The Result: It gives a single, perfectly accurate answer for the whole trend, even with small groups, without needing to run multiple separate tests.

Real-World Test: The Zomba Experiment

To prove their methods work, the authors applied them to a real experiment in Malawi called the Zomba Cash Transfer Program.

  • The Setup: They looked at schools where different percentages of girls were offered cash to stay in school.
  • The Question: Did the girls who weren't offered cash still benefit (or suffer) because their classmates were?
  • The Outcome: They ran their new tests. The results showed that, for the specific questions they asked, the data did not provide strong evidence that the spillover effects were happening in the way they hypothesized. Crucially, they showed that their new methods could handle the messy, small-group data of this real-world experiment much better than old methods could.

The Big Takeaway

This paper is like giving researchers a new, more reliable toolkit for studying how people influence each other in groups.

  • If you have a small group, they give you a method that is exact (no guessing).
  • If you care about averages, they give you a method that gets more accurate as you get more data.
  • If you have many levels of treatment, they give you a way to test the whole pattern at once.

They didn't invent new drugs or new teaching methods; they invented a better way to measure if those things work when people are influencing each other.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →