Learning What to Learn: Experimental Design when Combining Experimental with Observational Evidence
This paper proposes a unified framework for designing cost-constrained experiments that combine with potentially biased observational data by employing a minimax proportional-regret criterion to optimize the bias-variance trade-off without requiring pre-specified bias bounds.
Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Great Detective's Dilemma: When Clues Lie
Imagine you are a detective trying to solve a massive mystery: how a new policy, like giving money to poor families, will change an entire country's economy. You have two types of clues. The first type comes from a "Randomized Controlled Trial" (RCT). This is like a perfectly controlled science experiment where you give a small group of people a specific treatment and watch what happens in a lab. These clues are gold because they are trustworthy; you know exactly what caused the result. But there's a catch: these experiments are expensive and usually only happen in tiny, specific towns. They tell you what happens in that one village, but they can't tell you what happens if you roll the program out to the whole nation, where prices change and people react in complex ways.
The second type of clue comes from "Observational Evidence." This is like reading old case files or news reports about what happened in other places. These clues are cheap and cover huge areas, but they are messy. Maybe the people in those old reports were different, or maybe something else was going on that made the results look different than they really were. The big problem for scientists is: How do you mix these two types of clues? You want the trustworthiness of the experiment and the big-picture view of the old reports, but you don't know how much the old reports are lying. If you trust the old reports too much, your whole theory could collapse. If you ignore them, you waste a lot of useful information. This paper is about building a better map for detectives to decide exactly how much of their budget to spend on the expensive experiment versus how much to rely on the messy old files.
The Paper's Solution: The "Regret" Compass
This paper, titled "Learning What to Learn," proposes a new way to design experiments when you have to mix fresh, expensive data with old, potentially biased data. The authors, Aristotelis Epanomeritakis and Davide Viviano, argue that the old way of doing things—trying to minimize the "noise" or variance in your results—is dangerous if you ignore the possibility that your old data is biased. Instead, they introduce a new compass called Minimax Proportional Regret.
Think of "Regret" as the feeling you get when you look back and wish you had made a different choice. In this paper, the authors imagine a super-smart "Oracle" (a magical being) who knows the exact amount of bias in the old data. The Oracle picks the perfect experiment and the perfect way to mix the data to get the best answer. Since real researchers don't have a magic Oracle, they can't know the exact bias. So, the authors ask: "What if we pick a design that makes us regret our choice as little as possible, even in the worst-case scenario?"
The paper finds that the best strategy isn't just to pick the experiment that gives the most precise numbers. Instead, it's to find a "sweet spot" where you balance two things:
- Variance Regret: How "noisy" or shaky your results are because you didn't collect enough data.
- Bias Regret: How wrong your results could be because you trusted a biased old report too much.
The authors show that the optimal design is one where these two types of regret are equalized. It's like a tightrope walker: if you lean too far toward minimizing noise (collecting huge amounts of experimental data), you might fall off the cliff of bias (trusting a lie). If you lean too far toward minimizing bias (ignoring the noisy data), you might fall off the cliff of variance (your numbers are too shaky to be useful). The paper proves mathematically that the best design is the one where you are equally worried about both cliffs.
How It Works in the Real World
To prove their idea works, the authors tested it with a real-world example: a government in Kenya wanting to know how cash transfers affect school attendance. They had some old data from a famous program in Mexico (PROGRESA) that might not apply perfectly to Kenya. They had to decide whether to run a new experiment to measure the direct effect of cash, the effect of income, or the effect of wages.
Using their new "Regret" method, they found that the standard way of doing things (called "Neyman allocation") would have spent all their money on the cash-transfer experiments because those gave the cleanest numbers. However, this ignored the risk that the old Mexican data was biased. The new method suggested a different path: it recommended running a job program experiment at every sample size considered, rather than the cash-transfer experiments. The standard methods only switched to the job program at very large sample sizes, but the new method prioritized it immediately to avoid the risk of bias.
In their simulations, the new method showed that while the old methods might look great if the old data was perfect, they fell apart if the data was even slightly wrong. The new "Regret-optimal" design stayed steady. It accepted a small increase in the "noise" of the results to gain a massive reduction in the risk of being wrong. For example, at a sample size of 3,700 people (which is typical for these kinds of studies), the old methods had a "regret" score of over 4 or 5, meaning they were far from the best possible answer. The new method kept the regret score down to around 1.3, meaning it was much closer to the truth, even when the old data was biased.
What the Paper Says and Doesn't Say
The paper is very clear about what it does and doesn't do. It does not claim that we should stop doing experiments or that we should stop using old data. It does not say that we can magically know how biased the old data is. In fact, the whole point is that we don't know.
The authors explicitly argue against the idea of just trying to make the experimental numbers as precise as possible (minimizing variance) without thinking about bias. They show that if you do that, you might end up with a very precise answer to the wrong question. They also argue against being so scared of bias that you ignore the old data entirely; that wastes resources and leaves you with shaky results.
The findings are based on mathematical proofs and computer simulations. The authors show that their method works for a wide class of common statistical tools (like GMM and minimum distance estimators) and even for scenarios where a "Bayesian audience" (a group of people with different beliefs) tries to interpret the data. They prove that their "Regret" approach is a robust way to handle the unknown. They do not claim to have solved every possible problem in economics, but they provide a specific, mathematically sound tool for the very common problem of mixing clean experiments with messy real-world data.
The Takeaway
In the end, this paper gives researchers a new rulebook for spending their limited budgets. It tells them: "Don't just look for the shiniest, most precise data. Look for the design that keeps you safe from the biggest mistakes, even if you don't know exactly what those mistakes are." It's a lesson in humility: by admitting we don't know everything about the past, we can design better experiments for the future. The result is a more honest, more reliable way to answer big policy questions, ensuring that when governments decide to spend money on programs, they are basing those decisions on a foundation that is both precise and resilient.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.