← Latest papers
📊 statistics

A note on the optimum allocation of resources to follow up unit nonrespondents in probability

This paper proposes a method for optimally allocating nonresponse follow-up resources in probability surveys to minimize the mean squared error of estimators, rather than simply maximizing response rates, and demonstrates its application using data from an Australian agricultural survey.

Original authors: Su-Ming Tam, Anders Holmberg, Summer Wang

Published 2026-07-21
📖 8 min read🧠 Deep dive

Original authors: Su-Ming Tam, Anders Holmberg, Summer Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery about the entire population of a city, but you can only interview a small, random group of people. This is the world of probability surveys, the statistical tool governments and researchers use to understand everything from how many sheep are in a field to how people feel about their jobs. The big problem? Not everyone answers the door. Some people ignore the detective, some hang up the phone, and some simply refuse to talk. This is called nonresponse.

For a long time, the standard rule for these detectives was simple: "Knock on every single door that didn't answer until someone opens it." The goal was to get the highest possible response rate—the percentage of people who said "yes." The logic was that if you talk to more people, your picture of the city must be clearer. However, statisticians have realized that a high response rate doesn't always mean a high-quality picture. You might knock on a thousand doors just to find that the people who finally answered are all from the same neighborhood, leaving you with a biased view of the whole city. The real goal isn't just to get more "yes" answers; it's to get the most accurate answer possible with the money you have. This paper asks a crucial question: If you have a limited budget for knocking on doors, how should you spend it to get the best possible result, rather than just the loudest one?


The Detective's Dilemma: Why Knocking on Every Door Might Be a Waste of Time

In the world of official statistics, like those collected by the Australian Bureau of Statistics, there is a classic tug-of-war between cost and quality. Survey statisticians have a budget for "Nonresponse Follow-up" (NFU)—the money and time spent chasing down people who didn't answer the first time. The old-school strategy was to use this budget to follow up on every single non-respondent, hoping to turn them into respondents. The thinking was: "More respondents equals better data."

But the authors of this paper, SM Tam, A Holmberg, and S Wang, argue that this is like a detective spending all their money trying to interview the same uncooperative witness over and over again, just to say they tried, while ignoring other leads that could solve the case faster. They point out that response rate is a poor indicator of data quality. Sometimes, increasing the response rate actually makes the data worse by introducing more bias.

Instead of chasing a high response rate, the authors propose a new strategy: Minimize the Mean Squared Error (MSE). In plain English, MSE is a measure of how far off your guess is from the true answer. The goal is to spend your follow-up money in a way that makes your final estimate as close to the truth as possible, not just as loud as possible.

The New Strategy: The "Response Homogeneity Group" Map

To solve this, the authors suggest a method that sounds a bit like a video game strategy guide. They divide the non-respondents into groups called Response Homogeneity Groups (RHGs). Think of these groups as neighborhoods where everyone has a similar "willingness to talk" score, known as a propensity score.

  • Group A: People who are very likely to answer if you just ask nicely (High Propensity).
  • Group B: People who are very stubborn and unlikely to answer no matter what (Low Propensity).

The paper argues that you shouldn't treat everyone the same. If you have a limited budget, you shouldn't waste it trying to convince the super-stubborn Group B members to talk after you've already knocked five times. Instead, you should focus your energy on the groups where a few extra knocks will actually bring in new, useful information.

The authors use a mathematical framework called quasi-randomisation and inverse propensity weighting. Imagine you are taking a photo of a crowd. If you only photograph the people standing in the front row, your photo is skewed. But if you know exactly how many people are in the back row and how likely they were to step forward, you can mathematically "stretch" the photo to make it look like you saw everyone. This paper suggests using that math to decide who deserves the extra effort of a follow-up visit.

The Sheep Farm Experiment

To test their idea, the authors ran a simulation using real data from the 2018/19 Rural Environment and Agricultural Commodities Survey (REACS) in Australia. They were trying to estimate the number of sheep on farms.

Here is how they set up the experiment:

  1. The Data: They looked at 4,696 farms. 3,525 were continuing farms (they had been surveyed before), and 1,171 were new farms.
  2. The Prediction: They used a computer algorithm called Random Forest (a type of machine learning that acts like a team of decision-making trees) to predict how likely each farm was to respond. For the new farms, they used a "k-Nearest Neighbors" method, finding the most similar old farms to guess the new ones' behavior.
  3. The Groups: They sorted all the non-responding farms into 10 groups (RHGs) based on their likelihood to respond, ranging from "Very Unlikely" (0% to 10% chance) to "Very Likely" (90% to 100% chance).
  4. The Budget: They set a hypothetical follow-up budget of $30,000.

The Results: Spending Smarter, Not Harder

The authors used a computer solver (like a super-smart calculator) to figure out the perfect number of follow-up visits for each group to minimize the error in their sheep count, without blowing the budget.

The Old Way (Common Practice):
If you just knock on every door the same number of times, you might end up spending too much time on the super-stubborn farms (where you get no results) and not enough on the "maybe" farms. In their simulation, this "one-size-fits-all" approach cost $30,000 and resulted in a 99% response rate. However, the error in their sheep count was still relatively high.

The New Way (Optimal Allocation):
The authors' method suggested a very different plan:

  • For the "Very Unlikely" groups: Knock 6 times on the first group and 5 times on the second. These groups have high variance in their data, meaning they contribute the most to the overall uncertainty of the estimate. By allocating more visits here, the strategy targets the specific groups that will reduce the error the most.
  • For the "Very Likely" groups: Only knock 1 time. These people are already likely to answer; you don't need to waste money knocking 6 times on them.
  • The Cost: This smart strategy only cost $19,640 (about two-thirds of the budget).
  • The Result: Even though they spent less money, they reduced the nonresponse variance (the error caused by missing people) from 57.77 billion units down to 38.21 billion units.

The Big Surprise:
The "Common Practice" got a 99% response rate, while the "Optimal Strategy" only got a 97% response rate.

  • Common Practice: 99% response rate, higher error, cost $30,000.
  • Optimal Strategy: 97% response rate, lower error, cost $19,640.

The paper shows that the 99% response rate was a "false victory." It looked impressive, but it cost more money and actually gave a less accurate picture of the sheep population. The optimal strategy saved money and gave a better answer, even with fewer people saying "yes." Specifically, compared to the standard "knock on every door" method, the optimal allocation reduced the nonresponse variance by about 3.9%. (Note: The 10.4% reduction mentioned in the paper refers to the reduction in Root Mean Squared Error compared to doing no follow-up at all, not compared to the standard method).

What This Means for the Future

The authors conclude that chasing a high response rate is a trap. It's like trying to win a game by just counting how many times you yell "Check!" instead of actually checking if you have a winning hand.

Their findings suggest that survey statisticians should stop trying to convert every non-respondent. Instead, they should use data to identify which groups need a little extra push and which groups are already doing fine. By allocating resources based on where they will reduce the Mean Squared Error the most, agencies can get better data for less money.

In the specific case of the sheep survey, the optimal allocation reduced the error significantly compared to doing nothing, and by a smaller but meaningful margin compared to the standard "knock on every door" method, all while saving thousands of dollars. The paper doesn't claim this is a magic bullet for every problem, but it provides a clear, mathematically sound way to make survey follow-ups smarter, proving that sometimes, knowing when to stop knocking is just as important as knowing when to start.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →