← Latest papers
📊 statistics

Weighting a Census as a Non-Probability Sample: A Doubly Robust Framework for Correcting Differential Undercoverage in Uruguay's 2023 Census

This paper proposes a doubly robust weighting framework that combines response-propensity modeling with demographic calibration to correct for non-random undercoverage in Uruguay's 2023 Census, thereby producing unbiased population estimates and social indicators despite the limitations of available administrative records.

Original authors: Ferreira Juan Pablo, Goyeneche Juan Jose

Published 2026-06-05
📖 5 min read🧠 Deep dive

Original authors: Ferreira Juan Pablo, Goyeneche Juan Jose

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the 2023 Uruguayan Census as a massive attempt to take a photo of the entire country's population. Ideally, this photo would include everyone, perfectly capturing who they are, where they live, and what their lives are like.

However, the photo came out blurry in specific spots. About 10% of the people were missing from the picture. Worse, they weren't missing randomly; the camera missed the people who were struggling the most—the poor, the rural residents, and young adults. If you just counted the people who were in the photo, you would think the country was wealthier, older, and more urban than it actually is.

The statisticians at Uruguay's National Institute of Statistics (INE) had to fix this "blurry photo" without being able to go back and take the picture again. Here is how they did it, using a method they call Doubly Robust Weighting.

The Problem: Why Just "Adding" People Didn't Work

First, the government tried to fix the missing numbers by looking at their "digital footprints"—administrative records like tax files, health records, and utility bills. They found about 350,000 missing people this way.

Think of this like trying to fix a puzzle by finding loose pieces in a different box. While they found the number of missing people, these pieces were incomplete. The digital records told them the missing people's names, ages, and general location, but they didn't know if those people had a roof over their heads, how many kids they had, or if they were living in a crowded apartment.

If they just pasted these incomplete pieces into the puzzle, the picture would still be distorted. They needed a way to guess the missing details based on the people who were successfully photographed.

The Solution: The "Doubly Robust" Safety Net

The researchers treated the successfully photographed households as a "non-probability sample." In plain English, this means they admitted: "We didn't pick these people randomly; the people who let us in were easier to find. We need to adjust for that."

To fix this, they built a two-part safety net called a Doubly Robust framework. Imagine you are trying to guess the average height of a crowd, but you only see the people standing near the front.

  1. Model A (The "Who is Missing?" Guess): They looked at the neighborhoods (called "segments") and asked, "How hard was it to find people here?" They used a clever trick: they looked at how many people in that neighborhood successfully filled out the census online. If a neighborhood had a low online completion rate, they assumed it was also hard to find people there in person. This helped them guess who was missing.
  2. Model B (The "What Do They Look Like?" Guess): They used the known facts about the country (total men, total women, total people in each city) to create a "target profile." They asked, "If we had the whole country, what would the age and gender mix look like?"

The "Doubly Robust" Magic:
The beauty of this method is that it has a backup plan.

  • If Model A is perfect but Model B is slightly off, the result is still accurate.
  • If Model B is perfect but Model A is slightly off, the result is still accurate.
  • It only fails if both guesses are wrong.

This is like having two different maps to find a treasure. If one map is old and the other is new, you can still find the spot as long as at least one of them is right.

The Process: Stretching the Photo

Once they had these two models, they applied "weights" to the people who were counted.

  • If a person came from a neighborhood that was hard to reach (like a poor rural area), their "weight" was increased. Statistically, this means "One person in this photo represents 1.5 people in reality."
  • If a person came from an easy-to-reach area, their weight stayed close to 1.

They then adjusted these weights until the total number of men, women, and people in every city matched the official government totals. This ensured the final picture wasn't just a guess; it was mathematically forced to match the known reality of the country's demographics.

The Result: A Clearer Picture of Reality

When they applied this method, the picture changed significantly.

  • Before: The census looked like a country of older, wealthier people living in cities.
  • After: The weighted data revealed more young adults, more people living in rural areas, and more families struggling with housing issues (like overcrowding or poor materials).

For example, the percentage of people living in informal settlements (slums) jumped from 4.5% to 5.5%. This wasn't because more people moved there; it was because the original photo had missed the people living there. The new method "filled in the blanks" based on the patterns of the people who were found.

Why This Matters

This paper doesn't just say "we counted more people." It provides a new rulebook for how to count a country when the traditional method fails. It shows that when you can't get a perfect random sample, you can still get a reliable picture by combining:

  1. Smart guessing about who is missing (using local clues like internet usage).
  2. Hard facts about the total population (using government totals).
  3. A safety net that ensures you don't get it wrong just because one of your guesses was slightly off.

In short, they turned a flawed, biased snapshot into a reliable, representative portrait of the entire nation, ensuring that the most vulnerable people weren't invisible in the final statistics.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →