Randomization tests for model specification in causal inference under network interference
This paper proposes a novel design-based framework and a corresponding randomization-testing procedure to empirically assess the correct specification of exposure mappings in causal inference under network interference, providing theoretical guarantees and demonstrating effectiveness through simulations and a field experiment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of scientific experiments, researchers often rely on a simple, powerful idea: if you randomly assign a treatment to some people and not others, any difference in the outcome must be caused by that treatment. This logic works beautifully when people are isolated islands, unaffected by their neighbors. However, the real world is rarely so quiet. People are connected in complex webs of friendship, family, and community. When one person changes their behavior, it often ripples out to influence those around them. This phenomenon, known as a spillover effect, breaks the standard rules of experimentation. If a student adopts a new habit because their friend did, it becomes impossible to tell if the habit came from the original treatment or from the friend's influence. To make sense of this, scientists use a tool called an exposure map. This is a simplified way of describing how a person is exposed to a treatment, not just by their own status, but by the status of the people they are connected to. It reduces a chaotic web of influences into a manageable number, allowing researchers to estimate the true effect of an intervention.
The problem is that these maps are usually guesses. A researcher might assume that only immediate friends matter, or that the influence fades quickly with distance. If the guess is wrong, the entire experiment can lead to misleading conclusions. Until now, there has been no reliable way to check if a chosen map actually fits the reality of the situation. A new study by Supriya Tiwari and Pallavi Basu from the Indian School of Business addresses this gap. They have developed a new method to test whether the chosen exposure map is correct, using the very randomness of the experiment itself as a measuring stick. Instead of assuming the map is right, their approach asks the data to prove it.
The researchers built their method on the logic of randomization tests, a technique that has long been a gold standard in statistics. In a typical experiment, the treatment is assigned by chance. The authors realized that if their exposure map is truly capturing how the treatment spreads, the leftover errors in their predictions should be random and unrelated to the treatment assignment. But if the map is wrong, those errors will show a pattern, revealing that the model missed a crucial piece of the puzzle. To find these patterns, they looked at the connections between people. They checked if a person's treatment status was correlated with the prediction errors of their neighbors. If a treated person's neighbor had a large error in the prediction, it suggested the model failed to account for the influence flowing between them. By measuring this correlation across the entire network, they created a test statistic that could signal a misspecified model.
To ensure their method was sound, the team first proved mathematically that it works under ideal conditions, where the true underlying patterns are known. They then moved to simulations, creating thousands of fake networks and experiments on computers to see how the test performed in practice. They tested scenarios where the exposure map was perfect, and scenarios where it was deliberately flawed, such as ignoring the influence of second-degree friends or misjudging the strength of the spillover. The results were clear: when the map was wrong, the test almost always caught it, rejecting the incorrect model with high confidence. When the map was right, the test correctly accepted it, keeping false alarms low. The study also explored what happens when the mathematical model used to make predictions is slightly off, showing that the method remains robust and does not easily break under minor imperfections.
The researchers then applied their new tool to a real-world dataset from a famous field experiment involving middle school students. In that study, researchers tried to reduce bullying by training a small group of influential students, known as social referents, to promote anti-conflict norms. The original study assumed that the influence of these referents was captured simply by knowing if a student had a treated friend. Using the new testing procedure, the authors re-examined this assumption. They ran the test thousands of times by reshuffling the treatment assignments to see what the data would look like if the assumption were true. The result was a p-value of 0.317, a number that indicates the data is consistent with the original assumption. In other words, the simple map used in the original study held up under scrutiny; there was no statistical evidence to suggest it was missing a deeper layer of influence. This does not prove the map is perfect, but it confirms that the researchers did not need to discard it based on the available evidence.
This work offers a crucial new step for scientists studying interconnected systems. It moves the field from blindly trusting assumptions to actively verifying them. By providing a way to check if the chosen model of influence is correct, the method helps researchers avoid drawing false conclusions about how interventions spread through a population. It does not solve the problem of interference, nor does it guarantee that every model will be perfect. Instead, it provides a rigorous check, a way to say with confidence whether the map being used is a faithful representation of the territory. In a world where human behavior is deeply entangled, having a tool to verify our understanding of those connections is a significant advance, ensuring that the lessons learned from experiments are as reliable as the science that produces them.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.