← Latest papers
📊 statistics

Data integration of non-probability and probability samples with deterministic predictive mean matching

This paper proposes and validates deterministic predictive mean matching mass imputation estimators for integrating probability and non-probability samples, establishing their theoretical consistency and variance properties under both model specification and misspecification while demonstrating their effectiveness through simulations and an empirical application to job vacancy data.

Original authors: Aniela Czerniawska, Piotr Chlebicki, Łukasz Chrostowski, Maciej Beręsewicz

Published 2026-07-16
📖 5 min read🧠 Deep dive

Original authors: Aniela Czerniawska, Piotr Chlebicki, Łukasz Chrostowski, Maciej Beręsewicz

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to figure out the average height of every student in a massive school. The most reliable way to do this is to pick a few students completely at random (a "probability sample") and measure them. Because the selection was random, you can mathematically guarantee that your average is close to the truth. But what if you don't have time or money to measure those random students? Instead, you have a giant list of volunteers who signed up online to join a "tall people club" (a "non-probability sample"). This list is huge, but it's biased: it's full of basketball players and models, so the average height here is way too high. You can't just use this list, or your school's average will be wrong.

This is a classic puzzle in the world of statistics, a field dedicated to making sense of data. The challenge is "data integration": how do you mix a small, perfectly fair list with a huge, messy, biased list to get a perfect answer? Usually, statisticians try to guess the "rules" that connect the two groups (like how height relates to age or gender) and use those rules to fix the biased list. But what if those rules are slightly wrong? Or what if the data is so complicated that the rules break down? This is where the paper by Czerniawska and her team steps in. They are tackling the problem of how to combine these two very different types of data lists to get a reliable answer, even when the math gets tricky.

The authors propose a clever new way to do this mixing called "Predictive Mean Matching" (PMM). Think of it like a high-tech game of "Find Your Twin." Imagine you have a student from your fair random list (let's call him Alex) whose height you don't know yet, but you do know his age and grade. You look at the huge, biased online list to find a student who looks just like Alex based on those details. Once you find that "twin," you borrow their height and give it to Alex. By doing this for every student in the random list, you create a complete, accurate picture of the whole school.

The paper explores two specific ways to play this "Find Your Twin" game. The first method, which they call PMM A, looks for twins based on what the computer predicts the height should be. If the computer thinks Alex should be 170cm, it finds someone in the biased list who the computer also predicts to be 170cm. The second method, PMM B, is a bit more direct: it looks for someone in the biased list whose actual, measured height matches what the computer predicts for Alex.

The researchers spent a lot of time proving that these methods actually work. They showed that if you use enough data, these methods will eventually give you the correct answer, even if your computer's prediction rules aren't perfect. Specifically, they proved that the PMM A method (the one using predicted values to find matches) is robust to model misspecification under certain conditions. This is a big deal because other popular methods (like simply finding the "nearest neighbor" based on raw data) can fail miserably if the data gets too complex or if the prediction rules are slightly off. The authors proved that their "Predictive Mean Matching" approach is much more robust; it doesn't break as easily when the math gets messy.

They also figured out how to measure how "sure" we can be about the answer. When you borrow data from a biased list, there's a little bit of extra uncertainty, like a hidden wobble in your scale. The authors developed a new formula to calculate this wobble and showed how to use computer simulations (called "bootstrapping") to measure it accurately. They tested their ideas with thousands of fake data scenarios in a computer lab. In these simulations, their new methods performed just as well as the best existing tools when the rules were perfect, but they were much better when the rules were wrong or the data was complicated.

Finally, the team tried their method on real-world data from Poland. They wanted to estimate how many job vacancies were specifically aimed at Ukrainian workers. They combined a small, official government survey (the fair list) with a massive database of job postings from public employment offices (the biased list). The results showed that their method could successfully merge these two very different sources to give a clear, reliable estimate, proving that their "Find Your Twin" strategy works in the real world, not just in computer simulations.

In short, this paper provides a sturdy, reliable toolkit for statisticians who need to mix fair but small data with huge but biased data. It offers a way to get accurate answers even when the math isn't perfect, ensuring that decisions based on this data are built on solid ground rather than shaky guesses.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →