Program Evaluation with Remotely Sensed Outcomes
This paper proposes a method for nonparametrically identifying causal effects in experiments and quasi-experiments by combining experimental data with observational data where low-cost, scalable remotely sensed variables (like satellite imagery) serve as post-outcome proxies for imperfectly measured economic outcomes.
Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Problem: The "Blind" Experiment
Imagine you are a scientist trying to test a new fertilizer to see if it makes corn grow taller. You have a perfect plan: you give the fertilizer to half your fields (the Experimental Group) and nothing to the other half.
However, there is a catch. Measuring the exact height of every single corn stalk is incredibly expensive and time-consuming. You can't afford to do it for the experimental fields. So, you are left with a "blind" experiment: you know which fields got the fertilizer, but you don't know how tall the corn actually grew.
The usual fix: Researchers often try to use a cheap, imperfect proxy. Maybe they look at the color of the corn from a satellite photo. They assume, "If the satellite photo looks green, the corn must be tall." They train a computer on a different set of fields where they did measure the height, then apply that computer to their blind experiment.
The paper's discovery: The authors say this common method is broken. It's like trying to guess the weight of a person by looking at their shadow. If the light source changes, the shadow changes, even if the person's weight stays the same. In their world, the "shadow" is the satellite image, and the "person" is the economic outcome (like poverty or crop burning).
The Core Idea: The "Post-Outcome" Clue
The authors introduce a crucial distinction: Is the clue a cause or a result?
- The Old Way (Surrogate): Treats the satellite image as a cause or a middleman. (e.g., "The fertilizer causes the satellite image to change, which causes the corn to grow.") This is wrong.
- The New Way (Post-Outcome): Treats the satellite image as a result. The corn grows (the outcome), and because it grew, the satellite image changes. The image is a "fingerprint" left behind by the outcome.
Think of it like a crime scene.
- The Outcome: The thief stole the jewelry.
- The Remotely Sensed Variable: The muddy footprints left on the floor.
- The Logic: The theft caused the footprints, not the other way around. If you see footprints, you know a theft happened.
The Solution: The "Two-Recipe" Kitchen
The authors propose a method to combine two different "kitchens" (datasets) to solve the problem without ever measuring the corn height in the experimental fields.
Kitchen A (The Experimental Sample):
- What you have: You know who got the fertilizer (Treatment) and you have the satellite photos (The Clue).
- What you lack: You don't know the actual corn height (The Outcome).
- What you learn: You learn how the fertilizer changes the photos.
Kitchen B (The Observational Sample):
- What you have: You have a different set of fields where you do know the corn height and you have the satellite photos.
- What you lack: You don't know if they got the fertilizer (or the fertilizer wasn't randomized).
- What you learn: You learn how the corn height changes the photos.
The Magic Trick:
The authors assume that the "camera" (the satellite) works the same way in both kitchens. If a field has tall corn, the satellite photo looks a specific way, whether you are in Kitchen A or Kitchen B. This is called Stability.
By combining the two kitchens, they can mathematically "cancel out" the camera's quirks. They use the relationship between the fertilizer and the photo (from Kitchen A) and the relationship between the height and the photo (from Kitchen B) to solve for the missing link: How much did the fertilizer actually help the corn grow?
Why the Old Methods Fail
The paper points out that many researchers are currently using a "Two-Step" method that is fundamentally flawed:
- Step 1: Train a computer to guess corn height based on photos using Kitchen B.
- Step 2: Use that computer to guess the height in Kitchen A and compare the groups.
The Flaw: This method suffers from "attenuation bias." It's like trying to hear a whisper through a wall. The computer guesses the height, but because the photos aren't perfect, the computer's guesses are "fuzzy." When you compare the fuzzy guesses, the difference between the groups looks smaller than it really is. The authors show this method often underestimates the true effect by nearly half (47% in one of their real-world examples).
They also tested a newer method called "Prediction-Powered Inference" (PPI). They found that PPI only works if the two kitchens are identical in every way (same people, same time, same background). But in the real world, Kitchen A and Kitchen B are usually different (different years, different places). When they differ, PPI breaks down.
The Real-World Tests
The authors tested their new method on three real-world scenarios:
- Forest Cover in Uganda: Did paying people to save trees actually stop deforestation?
- Poverty in India: Did a new digital payment system reduce village poverty?
- Crop Burning in India: Did paying farmers to stop burning crop residue actually work?
The Results:
- In the poverty study, their new method gave results almost identical to the "gold standard" (where they actually measured the poverty), even though they didn't use the direct measurements for the main calculation.
- In the crop burning study, the old "Two-Step" method said the program worked a little bit. Their new method said the program worked much better (almost twice as effective). The old method was hiding the true success of the program.
The Takeaway
If you want to measure the success of a program using cheap, remote data (like satellite photos or phone signals), don't just treat that data as a direct substitute for the real answer. Instead, treat it as a fingerprint left behind by the result.
By acknowledging that the "fingerprint" is caused by the result, and by carefully combining data from a controlled experiment with data from the real world, you can get accurate answers without spending a fortune on expensive surveys. And, crucially, you can do this even if your computer models for predicting the outcome are imperfect.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.