Transporting causal effects from a randomized trial without "transportability:" a case study of political advertising during U.S. elections
This paper proposes a transfer learning framework with sensitivity analysis to estimate causal effects in a target population (Georgia) from a randomized trial in a source population (five battleground states) without assuming full transportability, revealing that small deviations from this assumption lead to heterogeneous ad effects on voter turnout driven by racial composition.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: Trying to Predict the Future from the Past
Imagine you are a political strategist. You just ran a massive experiment in five specific states (Pennsylvania, Wisconsin, Michigan, North Carolina, and Arizona) to see if sending negative digital ads against a candidate would get more people to vote.
The Result: In those five states, the ads did absolutely nothing. Voter turnout didn't change.
The Question: Now, you want to know: "Will these same ads work in Georgia?"
Georgia is a crucial state, but it's different. It has a different mix of people (more Black voters, different economic backgrounds) and a unique political history. The problem is that the original experiment didn't happen in Georgia, and we can't just assume the results will be the same.
This paper is about building a mathematical "time machine" and "translation guide" to answer that question, even when we can't be 100% sure the two places are comparable.
The Problem: The "Perfect Match" Myth
Usually, statisticians try to move results from one place to another by saying, "If we just adjust for the differences we can see (like age, race, and gender), the results should hold."
The authors call this the "Transportability" assumption. It's like saying, "If I have a recipe for a cake that works in a dry kitchen, and I adjust for the humidity in a wet kitchen, the cake will taste the same."
The Catch: What if there are invisible ingredients? Maybe the people in Georgia are more tired, or the news cycle is different, or they have different health insurance coverage. These are things the original experiment didn't measure. If these "invisible ingredients" exist, the standard math breaks down. The authors call this a violation of transportability.
The Solution: A "What-If" Safety Net
Instead of pretending we know everything, the authors propose a new framework that admits, "We don't know everything, but let's see how much our answer changes if we assume there are some invisible differences."
They use a concept called Sensitivity Analysis. Think of it like testing a bridge.
- Standard approach: "The bridge holds if the wind is calm."
- This paper's approach: "The bridge holds if the wind is calm. But what if the wind blows at 10 mph? 20 mph? Let's calculate exactly how much the bridge sways at every speed so we know when it might break."
In their math, they use a "sensitivity knob" (called ).
- Turn the knob to 1.0: We assume Georgia is exactly like the other five states (no invisible differences).
- Turn the knob to 1.01: We assume there is a tiny, invisible difference between the states.
- Turn the knob to 0.99: We assume the invisible difference goes the other way.
The Two Tools They Built
To do this math, they created two different calculators:
The "Simple Regression" Calculator (The Workhorse):
- Analogy: This is like using a reliable, standard map. It's easy to understand, easy to use, and gives you a good answer for most people.
- How it works: It looks at the data from the five states, adjusts for the visible differences (like race and age), and then applies a "correction factor" based on how much we think the invisible differences might matter. They use a technique called bootstrapping (resampling the data thousands of times) to make sure the answer is stable.
- Recommendation: The authors say, "If you are a practitioner, start with this one."
The "Efficient Influence Function" Calculator (The Race Car):
- Analogy: This is a high-performance race car. It's faster and theoretically more precise, but it's much harder to drive and requires a pit crew to maintain.
- How it works: It uses complex math to get the most accurate answer possible, but it requires estimating many more hidden variables.
- Warning: The authors note that unlike some other advanced tools, this one isn't "doubly robust." If you get one part of the math wrong, the whole answer could be wrong.
The "Calibration" Trick: How to Set the Knob
The hardest part of this method is deciding how far to turn the "sensitivity knob." How big should the invisible difference be? 1%? 10%?
Usually, people just guess. The authors invented a clever, data-driven way to figure this out, which they call Calibration.
- The Analogy: Imagine you want to know if a new car engine works in a snowy climate, but you only tested it in a desert. You don't know how much snow matters.
- The Trick: Instead of guessing, you take your desert data, split it into two groups: "Rust Belt" states (cold, industrial) and "Sun Belt" states (warm, sunny).
- You pretend the "Sun Belt" data is the new target. You try to transport the "Rust Belt" results to the "Sun Belt" using your math.
- You keep turning the sensitivity knob until your math predicts the Sun Belt results exactly as well as if you had just run the experiment there.
- The Result: The amount you had to turn the knob to make the math work tells you: "This is the size of the invisible difference between Rust Belt and Sun Belt." Now you know that if Georgia is that different from the original states, you should turn the knob that far.
What They Found in Georgia
When they applied this to the 2020 election data for Georgia, the results were surprising and nuanced:
- If we assume everything is perfect (Transportability holds): The ads still do nothing. Zero effect.
- If we assume there are tiny, invisible differences (Sensitivity Analysis): The result changes dramatically.
- In counties with more White voters and fewer Black voters, the ads might actually increase turnout (a positive effect).
- In counties with more Latinx voters, the ads might decrease turnout (a negative effect).
The Big Takeaway:
The authors conclude that because the original effect was so close to zero (it was basically a "null" result), it is extremely fragile. Even the tiniest, unmeasured difference between the five states and Georgia can flip the result from "no effect" to "big positive effect" or "big negative effect."
They quote a famous statistician, Paul Rosenbaum, to summarize: "Small effects are sensitive to small biases." Just because an experiment shows no effect in one place doesn't mean there is no effect elsewhere; it might just mean the effect is hiding behind a tiny, invisible difference that we didn't measure.
Summary
This paper teaches us that when moving scientific results from one group to another, we shouldn't just pretend the groups are identical. Instead, we should use a "safety net" to test how much our conclusions change if there are hidden differences. In the case of political ads, a "no effect" result in one place might actually be a "huge effect" in another, depending on the invisible details of the population.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.