← Latest papers
📊 statistics

Doubly robust integration of nonprobability and probability survey data

This paper extends doubly robust estimators for integrating nonprobability and probability survey data to domain estimation and proposes efficient combined estimators that leverage outcome data from the probability survey, providing theoretical variance formulae and demonstrating their finite-sample efficiency through simulations.

Original authors: Shaun R Seaman, Tommy Nyberg, Anne M Presanis

Published 2026-06-11
📖 5 min read🧠 Deep dive

Original authors: Shaun R Seaman, Tommy Nyberg, Anne M Presanis

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you want to know the average height of everyone in a massive city. You have two different ways to get this information, but both have flaws.

The Two Data Sources

  1. The "Gold Standard" Survey (Sample A): This is like a professional census. The government carefully picks people from a complete list so that every type of person is represented fairly. However, this survey is expensive and slow. Sometimes, they only ask people, "What is your height?" (Outcome data). They might not ask about your shoe size or age (Covariates), or they might not have enough people to answer the height question perfectly.
  2. The "Convenience" Survey (Sample B): This is like a social media poll or a survey done at a busy train station. It's easy to get a lot of people to answer, and they answer both "What is your height?" and "What is your shoe size?" (Outcome and Covariates). But, because people chose to answer themselves, the group is biased. Maybe only tall people with big shoes signed up. You can't just take their average; it's skewed.

The Problem

Statisticians have a clever trick called Doubly Robust (DR) Estimation. It's like a safety net. It combines the "Convenience" survey (which has lots of details) with the "Gold Standard" survey (which is representative) to fix the bias.

  • It uses the "Gold Standard" to figure out who is missing from the "Convenience" survey.
  • It uses the "Convenience" survey to guess the missing heights based on shoe sizes.
  • The Safety Net: If your guess about shoe sizes is wrong, the method still works because of the "Gold Standard" data. If the "Gold Standard" data is messy, the method still works because of the shoe size guess. As long as one of these two guesses is right, the final answer is accurate.

The New Twist in This Paper

The authors noticed a gap. Usually, statisticians had to choose: Do I use the "Gold Standard" data alone, or do I use the "Doubly Robust" trick?

But what if the "Gold Standard" survey also asked people their heights? Now you have height data from both sources!

  • Estimator 1: The "Gold Standard" average (purely based on the official survey).
  • Estimator 2: The "Doubly Robust" average (mixing the official survey's structure with the convenience survey's details).

The paper asks: How do we mix these two answers together to get the absolute best result?

The Solution: The Perfect Blend

The authors propose a new way to blend these two estimates. Think of it like mixing two different batches of paint to get the perfect color.

  • If the "Gold Standard" batch is very consistent (low noise) but the "Doubly Robust" batch is very detailed, you lean more on the Gold Standard.
  • If the "Doubly Robust" batch is very sharp and the Gold Standard is a bit fuzzy, you lean more on the Doubly Robust.
  • The Secret Sauce: The paper provides a mathematical recipe to figure out exactly how much of each to mix. It calculates the "noise" (variance) in both answers and how much they agree with each other (covariance). By mixing them perfectly, you get a final answer that is more precise than either one could be alone.

Key Findings (The "So What?")

  1. It's a Safety Net: The method is "doubly robust." Even if the statistical models used to fix the bias are slightly wrong, the final combined answer is still reliable.
  2. Subgroups Matter: The paper also shows how to do this for specific groups (like "only people aged 20–30"), even if there aren't many people in that group in the surveys. It borrows strength from the whole group to make the small group's estimate better.
  3. The Efficiency Limit: The authors found that while mixing the two methods makes the answer better, it has a limit. You can never get more than twice the accuracy of the single best method you started with. If one method is already terrible, mixing it with a good one won't help much. But if both methods are decent and have different strengths, the mix is a huge win.
  4. It Works in Real Life: They ran computer simulations (creating fake cities and fake surveys) to prove their math works. They found that when the two data sources are roughly the same size, the "perfect blend" gives the biggest improvement.

In Summary

This paper is about taking two imperfect ways of measuring a population—one that is representative but maybe sparse, and one that is detailed but biased—and mathematically fusing them. If the "Gold Standard" survey also has the outcome data, the authors show you exactly how to combine it with the "Doubly Robust" method to get the most accurate, reliable average possible, without needing to know if your underlying guesses are perfect.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →