Toward design-based inference for data integration
This paper proposes a design-based framework for integrating non-probability and probability samples by treating the former as a certainty stratum, enabling the development of consistent regression estimators that remain unbiased without relying on unverifiable missing-at-random assumptions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to count the total number of apples in a massive orchard to report to the city council. You have two ways to get this data:
- The "Volunteer" Pile: A group of people who love apples have already picked a huge pile of apples from the orchard and brought them to the gate. You know exactly how many apples are in this pile because they counted them. However, you don't know which trees they picked from. Maybe they only picked from the big, easy-to-reach trees, or maybe they only picked the reddest ones. You can't be sure if their pile represents the whole orchard.
- The "Official" Survey: You have a strict, scientific plan to walk through the orchard, pick apples from specific trees using a random method, and count them. This is the gold standard for accuracy, but it's expensive and slow.
The Problem:
Traditionally, statisticians try to mix these two piles together. They assume the "Volunteer" pile is a fair representation of the whole orchard (a big assumption called "Missing at Random"). If that assumption is wrong—if the volunteers only picked red apples, for example—the final count will be wrong, and no amount of math can fix it.
The Paper's Solution: A "Two-Step" Strategy
The authors of this paper propose a clever, step-by-step way to combine these two sources without making risky guesses about the volunteers' habits.
Step 1: The "Certainty" Zone
First, you treat the "Volunteer" pile as a completed zone. You say, "Okay, we have counted every single apple in this specific group. We don't need to guess anything about them; we just accept them as a known fact."
Step 2: The "Complementary" Survey
Next, you send your official survey team out, but with a strict rule: They are forbidden from looking at the trees that are already in the Volunteer pile. They only survey the remaining trees that the volunteers missed.
Because the survey team is only looking at the "leftover" trees, and they are using a strict random method, their data is perfectly reliable. You now have two distinct, non-overlapping groups:
- Group A: The known, fully counted volunteer pile.
- Group B: The scientifically sampled remainder of the orchard.
The Two Ways to Combine the Data
Once you have these two groups, the paper suggests two ways to calculate the total:
- The "Separate" Method: You calculate the average apple size for the volunteers and the average for the survey team separately, then add them up. This is the safest method. It works even if the volunteers were biased (e.g., if they only picked red apples) because you aren't trying to force their data to look like the survey data.
- The "Combined" Method: You assume the volunteers and the survey team are basically the same type of people and mix their data together to get a single average. This is more efficient (gives a tighter estimate) only if the two groups are actually similar.
The "Pilot" Trick (Making the Survey Smarter)
Here is the clever part: Because you already have the huge "Volunteer" pile, you can use it as a pilot study to help your survey team work smarter.
Before the survey team goes out, you look at the volunteer pile to see how much the apple sizes vary. If you see that big trees have wildly different apple sizes, you can tell your survey team: "Hey, go check more trees that look like those big ones." This helps you design a better survey that gets the most accurate answer with the least amount of work. This is called optimal sampling.
Why This Matters (The Results)
The authors tested this idea with computer simulations and real data from Lithuanian government statistics (counting business sales and pharmacy drug sales).
- It's Robust: Even when the "Volunteer" pile was heavily biased (picking only specific types of apples), the new method stayed accurate. Old methods that tried to "fix" the bias mathematically failed miserably in these cases.
- It's Efficient: When the two groups were similar, mixing them together gave a slightly better answer. When they were very different, keeping them separate was safer.
- No Magic Assumptions: The method doesn't require you to believe that the volunteers picked apples randomly. It just treats them as a known block of data and does the hard math on the rest.
The Bottom Line
Think of this paper as a new rulebook for mixing "messy, unverified data" with "clean, scientific data." Instead of trying to force the messy data to fit a perfect model (which often fails), it isolates the messy data as a known block, uses it to help design a better survey for the rest of the world, and then combines the results in a way that is mathematically guaranteed to be accurate, no matter how weird the messy data was.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.