Disentangling spatial interference and spatial confounding biases in causal inference
This paper clarifies the definitions of spatial interference and spatial confounding using a Directed Acyclic Graph (DAG) framework to derive analytical bias expressions under general distributional settings, demonstrating how spatial weights, treatment distributions, and interference magnitude critically influence causal estimates while validating these theoretical findings through simulations and real-world data.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to figure out how much a specific fertilizer (the Treatment) helps a garden grow (the Outcome). You have a map of many different garden plots.
In a perfect world, if you put fertilizer on Plot A, only Plot A grows better. But in the real world, things are messy. This paper tackles two specific ways this messiness tricks you into drawing the wrong conclusions about your fertilizer.
Here is the breakdown of the paper's findings using simple analogies:
1. The Two Big Problems: "The Neighbor Effect" and "The Hidden Guest"
The authors say there are two main reasons your garden experiment might go wrong:
- Spatial Interference (The Neighbor Effect): Imagine your garden plots are right next to each other. If you water Plot A, the water might seep underground and accidentally water Plot B. So, Plot B grows better, not because of its own fertilizer, but because of Plot A's water. In the paper, this is called Spatial Interference. If you ignore this, you might think your fertilizer is working when it's actually just the "spillover" from a neighbor.
- Spatial Confounding (The Hidden Guest): Imagine there is a hidden factor, like a hidden underground spring, that affects both where you decide to put fertilizer and how well the plants grow. Maybe you only put fertilizer in the wet areas because the soil is soft there, and those wet areas naturally grow better plants. If you don't account for this "Hidden Guest" (the confounder), you will think the fertilizer is doing all the work, when really the water was the hero.
2. The "Hidden Guest" has two faces (Direct vs. Indirect)
The paper makes a crucial discovery: The "Hidden Guest" isn't just one thing. It comes in two flavors, and previous studies often treated them as the same.
- Direct Confounding: The hidden spring is under Plot A, affecting Plot A's fertilizer choice and Plot A's growth.
- Indirect Confounding: The hidden spring is under Plot A, but it flows underground to affect Plot B's growth and Plot B's fertilizer choice.
The Paper's Claim: These two are different! Just like a neighbor watering your garden (Interference) is different from a hidden spring affecting your soil (Confounding), a "Direct" hidden guest is different from an "Indirect" one. If you treat them as the same, your math breaks.
3. The "Shape" of the Data Matters
Most scientists assume that garden data looks like a perfect bell curve (Normal distribution). The authors say, "Wait a minute, real life isn't always a bell curve."
They tested what happens if your data looks like a Poisson distribution (think of counting things, like the number of raindrops or visits to a hospital) instead of a smooth curve.
- The Finding: The amount of error (bias) in your results changes depending on the "shape" of your data. If you assume a bell curve when you actually have "count" data, your calculation of the error will be wrong.
4. The "Map" You Choose Changes the Result
To measure how neighbors affect each other, you need a "Weight Matrix" (a map that decides who is a neighbor).
- The Paper's Claim: How you draw this map matters immensely.
- If you use Distance (e.g., "everyone within 5 miles is a neighbor"), the error gets smaller as you make the distance bigger.
- If you use K-Nearest Neighbors (e.g., "the 4 closest plots are neighbors"), the error actually gets worse as you add more neighbors.
- Analogy: It's like deciding who counts as your "friend." If you define a friend as "someone within 1 mile," you get one result. If you define a friend as "the 4 closest people regardless of distance," you get a totally different result. The paper shows that picking the wrong definition changes your final answer.
5. The Real-World Test: Manchester Weather
The authors didn't just do math on paper; they tested this on real weather data from Manchester, UK.
- The Setup: They looked at Air Temperature (Treatment) and Land Surface Temperature (Outcome). They knew Rainfall was a "Hidden Guest" (Confounder) because rain cools the air and wets the ground.
- The Result: When they ignored the "Neighbor Effect" (Interference) or mixed up the "Direct" and "Indirect" hidden guests, their estimates of how air temperature affects land temperature were wildly different and often wrong.
- The Winner: The only model that gave a reliable answer was the one that accounted for all three: The Neighbor Effect, the Direct Hidden Guest, and the Indirect Hidden Guest.
Summary: What Should You Take Away?
- Don't assume neighbors don't talk to each other. In spatial data, what happens next door often affects you. Ignoring this leads to bad conclusions.
- Don't lump all "hidden factors" together. A hidden factor affecting your immediate area is different from one affecting your neighbor's area. You have to separate them to get the right answer.
- Check your data's shape. If your data is "count" data (like Poisson) and not a smooth bell curve, the standard math for finding errors doesn't work.
- Be careful with your "neighbor map." The way you decide who is a neighbor changes your results.
The Bottom Line: If you are studying anything on a map (weather, disease, crime, crops) and you ignore how locations influence each other or mix up different types of hidden influences, your results will be biased. You might think a treatment works when it doesn't, or vice versa. To get the truth, you have to untangle all these threads at once.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.