When and How to Pilot: Design Rules for Two-Wave Experiments
This paper proposes a Conditional Minimax Regret (CMR) rule for two-wave experiments that optimally balances the robustness of balanced assignment with the efficiency of Neyman allocation by minimizing worst-case regret over a finite-sample confidence set, thereby avoiding the severe precision losses of adaptive methods in small pilots while capturing their gains in large ones.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery, but you have a limited supply of clues. You want to know which suspect is guilty, but you don't know who is hiding the most secrets. In the world of science, researchers run "experiments" to find answers, like testing if a new medicine works or if a new teaching method helps students learn. To do this, they split people into two groups: a "treatment" group that gets the new thing, and a "control" group that gets nothing (or the old thing). The goal is to compare the two groups to see the difference.
But here's the tricky part: not everyone reacts the same way. Some people are very consistent; others are all over the place. If the "treatment" group is full of unpredictable people, you need to test more of them to get a clear answer. If the "control" group is full of wild cards, you need more of them. The big question is: how do you know which group is the "wild card" group before you start your main investigation?
Scientists often run a tiny "pilot" test first—a small taste of the real thing—to get a hint. But this creates a dilemma. If you trust the pilot too much, you might get tricked by a fluke (like flipping a coin twice and getting heads both times, thinking the coin is rigged). If you ignore the pilot completely, you might miss a chance to get a much sharper answer. This paper is about finding the perfect balance between being too cautious and being too reckless when using those tiny pilot hints to design a big experiment.
The Great Experiment Balancing Act
Imagine you are a chef preparing a massive banquet for 1,000 guests. You have two main dishes: a spicy curry and a mild stew. You want to know which one people like better. But you also know that some people are very picky eaters (their taste varies wildly), while others are very consistent. If you serve the spicy curry to a group of picky eaters, you'll need to serve it to many of them to figure out the average opinion. If you serve the mild stew to a group of consistent eaters, you don't need as many.
The "perfect" plan would be to serve the curry to 800 people and the stew to 200, if the curry is the one with the picky eaters. This is called the Neyman allocation (named after a statistician who figured this out decades ago). It's the most efficient way to get an answer.
But here's the catch: you don't know who the picky eaters are yet! You have to decide the split before the banquet starts.
So, you decide to cook a tiny "pilot" meal first. You invite 30 people to taste-test the dishes. Crucially, each person only gets to try one dish. Some get the curry, and some get the stew. Based on how those specific groups reacted, you plan the big banquet.
The Old Ways:
- The "Don't Rock the Boat" Chef: This chef ignores the pilot entirely. They just split the 1,000 guests 50/50. It's safe. No matter what happens, they won't accidentally run out of curry for the picky eaters. But they might miss out on getting a super-precise answer if the pilot actually showed a big difference.
- The "Trust the Pilot" Chef: This chef looks at the 30 pilot tasters and says, "Hey, the curry group was super unpredictable! I'll serve curry to 900 people!" But what if the 30 pilot tasters just happened to be having a bad day? What if the curry is actually fine, but the pilot group was just noisy? This chef might end up serving curry to 900 people when they only needed 200, wasting resources and getting a worse answer than if they had just split it evenly. In the worst case, if the pilot group for one dish was perfectly calm (everyone liked it exactly the same), this chef might decide to serve zero people that dish, leaving them with no data at all.
The New Solution: The "Conditional Minimax Regret" (CMR) Rule
The author of this paper, Juan Yamin, proposes a new way to think about this. He calls it the Conditional Minimax Regret (CMR) rule.
Think of CMR as a very smart, slightly paranoid chef who uses the pilot data to draw a "safety zone" instead of a single guess.
- Instead of saying, "The curry is definitely unpredictable," the chef says, "Based on these 30 people, the curry could be unpredictable, but it could also be calm. Here is a range of possibilities."
- If the pilot data is weak (only 30 people), the "safety zone" is huge. The chef realizes, "I can't be sure yet," so they stick close to the safe 50/50 split.
- If the pilot data is strong (500 people), the "safety zone" shrinks. The chef can now say, "Okay, I'm pretty sure the curry is the wild card," and they shift the split toward the curry group.
The magic of CMR is that it never moves too far unless the evidence is strong enough to rule out the "bad luck" scenarios. It comes with a certificate—a little note that says, "Even if I'm wrong, here is the worst-case scenario, and it's not that bad."
What the Paper Found
The paper doesn't just talk about this; it proves it with math and tests it with simulations using real data from famous past experiments (like studies on deworming kids in Africa or testing resumes for job discrimination).
The Results:
- The "Trust the Pilot" Chef (Feasible Neyman) is dangerous with small pilots. When the pilot is small (like 30 people), this method often makes huge mistakes. In the simulations, it sometimes caused the experiment to be so inefficient that it was as if the researchers had thrown away half their data. In some cases, it even tried to assign zero people to a group that actually needed testing, which is a disaster.
- The "Don't Rock the Boat" Chef is safe but misses opportunities. They never make a huge mistake, but they also never get the extra precision that a good pilot could offer.
- The CMR Chef is the Goldilocks. When the pilot is small, CMR stays close to the safe 50/50 split, avoiding the huge disasters of the "Trust the Pilot" chef. But as the pilot gets bigger and the data gets clearer, CMR smoothly shifts toward the perfect split, capturing almost all the benefits of the "Trust the Pilot" method without the risks.
In the simulations, when the pilot was small (30 people), the "Trust the Pilot" method sometimes failed completely (infinite loss), while CMR stayed safe. As the pilot grew to 500 people, CMR caught up and performed almost as well as the perfect method, but with a guarantee that it wouldn't crash and burn if the pilot was just a fluke.
Why It Matters
This paper solves a very specific, annoying problem that scientists face every day: How much should I trust my small test run?
The answer is: Trust it just enough to move, but not so much that you fall off a cliff.
The paper shows that you don't have to choose between being a coward (ignoring the pilot) or a gambler (betting everything on the pilot). You can use a rule that says, "I will move my plan exactly as far as the evidence allows me to be sure I'm not making a mistake." It turns the scary uncertainty of small data into a manageable, safe step forward.
The authors tested this with four different real-world scenarios, and in every case, their new rule protected the researchers from the worst mistakes while still letting them get the best answers when the data was good. It's like having a seatbelt that only tightens when you actually hit a bump, keeping you safe without making the ride uncomfortable.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.