← Latest papers
📊 statistics

Coarsening-adjusted estimators for finite-population indicators

This paper proposes a design-based framework that addresses self-report coarsening in finite-population surveys by jointly modeling latent variables and reporting regimes via a survey-weighted pseudo-likelihood, generating posterior predictive replicates to apply standard estimators while explicitly propagating uncertainty through variance decomposition.

Original authors: Gaia Bertarelli, Aldo Gardini

Published 2026-07-28
📖 5 min read🧠 Deep dive

Original authors: Gaia Bertarelli, Aldo Gardini

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to count the exact number of jellybeans in a giant jar, but instead of letting people look inside, you ask them to guess. Most people, wanting to be quick or feeling unsure, won't say "47." They'll say "50." Others might say "40" or "60." In the world of statistics, this is called "coarsening" or "heaping." It happens all the time in surveys about health, money, or habits. People round up their income, estimate their weight, or guess how many cigarettes they smoke. The problem is, when statisticians try to calculate important numbers like the average or the percentage of people above a certain limit, these rounded guesses can trick the math. It's like trying to build a precise bridge using only rounded-off measurements; the structure might look okay from a distance, but up close, the beams don't quite line up.

This paper tackles a specific headache for scientists who run big surveys: how to fix these rounded numbers without throwing away the complex rules used to pick the people being surveyed. Usually, statisticians have two choices. They can either ignore the rounding (which leads to wrong answers) or use fancy computer models that assume the whole world follows a simple pattern (which often fails when real life is messy). The authors, Gaia Bertarelli and Aldo Gardini, propose a clever middle ground. They treat the rounded number you see as a "shadow" of a real, hidden number. They build a model to guess what that hidden number probably was, but they do it in a way that respects the survey's specific design. Think of it as a detective who doesn't just guess the suspect's height based on a blurry photo, but uses the photo, the lighting, and the known rules of the crime scene to reconstruct the most likely reality.

The Detective Work of Data

The paper introduces a new method to clean up these "fuzzy" survey answers. The authors treat the number a person reports (like "20 cigarettes") as a coarsened version of a secret, exact number (maybe 18 or 22) that they actually have in their head. They use a statistical trick called "multiple imputation," which is like running a simulation game thousands of times. In each game, the computer guesses a different "real" number for every person based on the rounded answer they gave and some other clues (like their age or gender).

Once the computer has generated these thousands of "reconstructed" datasets with the hidden numbers filled in, the statisticians apply their standard survey tools to each one. This is the magic part: by looking at how the results change across all these different versions, they can measure two things at once. First, how much uncertainty comes from the survey design (like picking a small group of people). Second, how much extra uncertainty comes from the fact that the original answers were rounded. It's like measuring the wobble of a table: some wobble comes from the uneven floor (the survey design), and some comes from the fact that the table legs are slightly bent (the rounding). The paper shows how to separate these two sources of error so officials know exactly how much to trust their numbers.

Testing the Theory with Real Habits

To see if their idea works, the authors tested it on real data from Italy's PASSI health survey. They looked at two very common habits: how many cigarettes people smoke daily and how many minutes they spend doing moderate exercise. These are classic cases of heaping. People tend to say they smoke in multiples of 5 or 10 (like 10, 20, or 30) and exercise in multiples of 30 or 60 minutes.

The researchers found that ignoring the rounding leads to distorted pictures of reality. For example, if you just count the people who say "20 cigarettes," you might think the number of "heavy smokers" is higher or lower than it really is, depending on how the rounding happened. Their new method, which they call "coarsening-adjusted," successfully smoothed out these artificial spikes. They showed that while the average number of cigarettes might not change much after fixing the rounding, the number of people crossing the "heavy smoker" threshold (20 cigarettes) could shift significantly. This is crucial because public health policies often rely on these specific cut-off points.

The paper also ran computer simulations to see how their method holds up when the assumptions aren't perfect. The results suggest that their approach is robust; even if the model guesses the "shape" of the hidden data slightly wrong, the final estimates for things like averages, medians, and thresholds remain much more accurate than the old, naive way of just using the rounded numbers.

Why It Matters

The main takeaway is that rounding isn't just a small annoyance; it can systematically skew the indicators that governments and health agencies use to make decisions. By using this new framework, statisticians can generate "posterior predictive replicates"—essentially, a cloud of possible true values for every respondent—and then apply standard survey formulas to them. This allows them to formally account for the extra uncertainty caused by rounding.

The authors demonstrate that different types of data react differently to this problem. A simple average might be relatively stable, but a threshold-based indicator (like "what percentage of people smoke more than 20 cigarettes?") can be highly sensitive to where the rounding happens. If a policy target is set at exactly 20 cigarettes, and many people round their 19 or 21 down or up to 20, the policy might look like it's failing or succeeding purely because of how people report their habits. The paper doesn't claim to have solved every problem in statistics, but it offers a powerful, flexible toolkit that bridges the gap between complex modeling and the practical needs of official statistics, ensuring that the numbers used to guide public health are as clear and accurate as possible.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →