← Latest papers
📊 statistics

Experimental Design under Network Interference

This paper proposes a statistical framework for optimizing two-wave experimental designs under network interference by leveraging pilot study data to minimize estimator variance, while formally characterizing the trade-offs involved in pilot size and providing theoretical guarantees for inference and regret.

Original authors: Davide Viviano

Published 2026-06-23
📖 5 min read🧠 Deep dive

Original authors: Davide Viviano

Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a city planner trying to figure out if a new park improves the happiness of the people living nearby. You want to build the park in the best possible spot to get the clearest answer. But there's a catch: people in this city are very social. If you build a park for one family, their neighbors might also feel happier just because they can see the greenery, even if they don't live right next to it. This is called "spillover" or "interference."

In traditional experiments, you might just pick a few random houses and give them the park, then compare them to houses without one. But in a connected city, this is messy. If you pick a house, its neighbors are affected, and if you pick a neighbor, they affect the first house. It's like trying to measure the splash of a single pebble in a pond that's already full of ripples from other pebbles.

This paper, by Davide Viviano, proposes a clever two-step strategy to solve this mess and get the most precise answer possible. Think of it as a "test run" followed by the "main event."

The Two-Step Strategy

Step 1: The "Scout Team" (The Pilot Study)

Before you build the park for everyone, you send out a small scout team to a tiny, isolated part of the city.

  • The Goal: You want to learn how much happiness varies from house to house and how much neighbors influence each other.
  • The Trick: You must pick a scout team that is isolated. They shouldn't have neighbors who are not on the scout team. If a scout's neighbor is left out, that neighbor's mood might change based on the scout's park, and that "contamination" ruins your data.
  • The Math: The paper uses a special computer algorithm (like a smart maze solver) to find the perfect group of houses that are clustered together but cut off from the rest of the city. This ensures the data they collect is "clean."

Step 2: The "Main Event" (The Real Experiment)

Once the scouts return with their data, you know exactly how much happiness varies and how strong the neighbor effects are. Now, you can design the main experiment perfectly.

  • The Goal: You want to pick the best houses for the main park and decide exactly who gets the park and who doesn't, specifically to minimize the "noise" (statistical variance) in your final result.
  • The Optimization: Instead of picking houses randomly, you use the scout's data to run a complex calculation. It's like a chess player who looks ahead to see which moves will lead to the clearest win. You might pick houses that are very different from each other or group them in a specific way to cancel out the noise.
  • The Rule: The houses you pick for the main event must be far away from the scout team (and their neighbors) so that the scout's "contamination" doesn't mess up the main results.

The Big Trade-Off

The paper highlights a delicate balancing act, like walking a tightrope:

  • Too small of a scout team: You don't have enough data to know how the neighbors influence each other. Your main experiment might be poorly designed because you're guessing.
  • Too large of a scout team: You learn a lot, but you have to exclude a huge chunk of the city from the main experiment (because you can't use anyone near the scouts). This leaves you with fewer people to study in the main event, which also makes your results less precise.

The paper provides a mathematical "sweet spot" (a rule of thumb) for how big the scout team should be relative to the main team to get the best possible result.

Why This Matters

The paper shows that this two-step method is much better than the old ways of doing experiments (like just picking random clusters or saturating groups).

  • Real-World Proof: The author tested this on real data from villages in Africa and simulated city networks. The results showed that this method could reduce the "error" in the results by up to 40% compared to other methods.
  • Handling Chaos: It works especially well when the data is messy (some people are naturally happier than others, and neighbors affect each other differently).

The Bottom Line

If you want to know the true effect of a treatment (like a new policy, a drug, or a community program) in a world where people influence each other, don't just guess and check.

  1. Send a small, isolated scout team to map the terrain.
  2. Use that map to strategically place your main experiment, avoiding the scout's neighborhood.
  3. This gives you the clearest, most precise answer possible, saving you time and money by needing fewer people to get a reliable result.

The paper is essentially a guidebook for running the smartest possible experiment in a connected world, ensuring that the "ripples" from one person don't drown out the signal you're trying to hear.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →