Power-Optimal Covariate Adjustment for Switchback Experiments
This paper proposes a power-optimal CUPAC methodology for switchback experiments with unequal cluster sizes that improves statistical power by specifically balancing the prediction of noise between and within randomization units, rather than merely targeting overall predictive accuracy.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of online platforms, where millions of transactions happen every hour, companies rely on experiments to decide if a new feature actually works. Imagine a food delivery app testing a new way to route drivers. To know if the change helps, the company cannot simply try it on half its users and ignore the other half; the users are too interconnected, and the timing of their orders matters. Instead, researchers use a method called a switchback experiment. In this setup, the entire system switches between the old way and the new way in short bursts, like a lighthouse beam sweeping across the water. Every hour, or every few minutes, the whole network flips to the new method, then flips back. This creates a series of time blocks where the entire system is treated as a single unit.
The goal of these experiments is to measure the difference in outcomes, such as how long it takes to deliver food, with as much precision as possible. To get a clear signal from the noise, data scientists often use a technique called covariate adjustment. Think of this as using a weather forecast to predict today's temperature. If you know the temperature yesterday and the season, you can predict today's temperature much better than if you just guessed. In an experiment, scientists use data from before the test started to predict what would have happened without the new feature. By comparing the actual results to these predictions, they can strip away the random fluctuations and see the true effect of the change. This process is standard practice, but it assumes that every piece of data is equally important.
A new study by Sergei Pankratev at DoorDash challenges this assumption for switchback experiments. The research reveals that in these specific types of tests, the standard way of using past data to clean up results is actually missing the most important part of the story. In a switchback experiment, the data comes in chunks: a specific group of drivers and customers during a specific hour. These chunks, or cells, vary wildly in size. Some hours are busy with thousands of orders; others are quiet with only a few. The standard method treats every single order as an equal data point, trying to predict the outcome for each individual order. However, because the experiment switches the entire system at once, the real source of uncertainty is not the individual orders, but the differences between these time chunks. The standard method spends its effort trying to predict the noise of individual orders, which the experiment design already averages out, while ignoring the larger swings between the time blocks that actually determine whether the experiment succeeds or fails.
Pankratev developed a new approach that fixes this misalignment. Instead of training the prediction model to be accurate for every single order, the new method trains the model to be accurate for the average outcome of each time block. It forces the computer to pay attention to the big picture: how the system behaves during a busy hour versus a quiet one. The study shows that this shift in focus is not just a minor tweak; it is a fundamental reweighting of what the model learns. The researchers found that in situations where the individual orders are highly unpredictable, the standard method fails to reduce the uncertainty of the experiment. It leaves the results fuzzy and makes it hard to tell if a change is real. The new method, by contrast, sharpens the focus on the time blocks, significantly reducing the uncertainty and making the experiment much more powerful.
To prove this, the researchers ran thousands of simulated experiments on a computer. They created a digital world with 200 different locations and 24 hours of activity, generating 4,800 distinct time blocks. In these simulations, they tested two different ways of training the prediction model. One group used the traditional method, which treats every order equally. The other group used the new method, which prioritizes the accuracy of the time-block averages. They also varied the complexity of the computer models, making them simple with few decision points and complex with many. The results were clear and consistent. When the individual orders were chaotic and unpredictable, the traditional method struggled. Even when the researchers made the computer models much larger and more powerful, the traditional method did not get much better. It was like giving a better telescope to someone looking at the wrong part of the sky; the extra power was wasted because the target was misaligned.
The new method, however, thrived in these chaotic conditions. By focusing on the time blocks, the model learned to predict the swings in the system that mattered most. As the models grew larger, the new method became even more effective, steadily reducing the uncertainty of the experiment. The study showed that this improvement translated directly into a higher chance of detecting a real effect. In the most difficult scenarios, where the noise was highest, the new method increased the statistical power of the experiment by 26 to 35 percentage points compared to the old way. This means that with the same amount of data, a company could be much more confident in its results, or it could achieve the same confidence with far fewer hours of testing.
The research also addressed a second part of the process: how the final numbers are calculated after the experiment is over. The study found that using the new training method is not enough on its own. If the model is trained to focus on time blocks, but the final calculation still treats every order as equal, the benefits are lost. The researchers showed that the calculation step must also be adjusted to match the training. When both the training and the calculation are aligned to focus on the time blocks, the experiment reaches its full potential. If they are mismatched, the gains disappear. This two-step alignment ensures that the effort put into training the model is fully realized in the final answer.
The findings suggest that for companies running these types of experiments, the way they prepare their data is just as important as the experiment design itself. The study does not claim that the old method is useless; in situations where the system is very stable and the time blocks are similar, the old method works fine. But in the messy, real world of online platforms where some hours are packed and others are empty, the old method leaves a lot of value on the table. The new approach offers a way to extract more clarity from the same data, turning a noisy signal into a clear answer. It is a reminder that in science, the tools we use to measure the world must be tuned to the specific nature of the world we are measuring. By reweighting the focus from the individual to the group, the researchers have provided a more precise lens for seeing how changes truly affect the systems we rely on every day.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.