← Latest papers
📊 statistics

Learning from Biased and Costly Data Sources: Minimax-optimal Data Collection under a Budget

This paper proposes a minimax-optimal data collection strategy under a fixed budget that maximizes effective sample size by optimizing the aggregated source distribution to minimize χ2\chi^2-divergence from the target, thereby outperforming naive sampling and standard estimators for estimating population and group-conditional means.

Original authors: Michael O. Harding, Vikas Singh, Kirthevasan Kandasamy

Published 2026-06-17
📖 5 min read🧠 Deep dive

Original authors: Michael O. Harding, Vikas Singh, Kirthevasan Kandasamy

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a case about the entire population of a city. You need to gather clues (data) to figure out the average height of everyone, or perhaps how different groups (like children vs. adults) differ in height.

However, you have a strict budget. You can't just go out and ask everyone; you have to buy your information from specific "sources," like a local park, a high school, or a retirement home.

Here is the catch:

  1. Different Prices: Asking a question at the park costs $1, but asking at the high school costs $5.
  2. Different Mixes: The park is full of kids (cheap to ask), but the city is mostly adults. The high school is mostly adults (expensive to ask), but the city also has many seniors.
  3. The Trap: If you just buy the cheapest data (the park), you get 1,000 answers, but they are all kids. If you try to "match" the city's demographics exactly, you might spend all your money on the expensive high school and only get 200 answers.

The Paper's Big Idea:
The authors, Michael Harding, Vikas Singh, and Kirthevasan Kandasamy, argue that the old way of thinking is wrong. You shouldn't just try to buy a "perfectly balanced" sample, nor should you just buy the cheapest sample.

Instead, they propose a new strategy based on something they call "Effective Sample Size."

The "Effective Sample Size" Analogy

Imagine you are baking a cake. You need flour, sugar, and eggs.

  • Naive Approach 1: You buy 1,000 bags of flour because it's cheap. You have a mountain of flour, but no cake.
  • Naive Approach 2: You try to buy exactly the right ratio of ingredients to match a recipe perfectly, but because eggs are expensive, you can only afford to buy 10 eggs and 10 cups of flour. You have a tiny cake.
  • The Authors' Approach: They realize that even if your ingredients are "biased" (you have too much flour), you can still bake a great cake if you have a smart baker (a special math formula) who knows how to weigh the ingredients correctly.

The goal isn't to buy the "perfect" mix of ingredients. The goal is to buy the maximum amount of ingredients that, once the smart baker adjusts them, gives you the most accurate cake.

They found that the "quality" of your data depends on a specific math formula involving χ2\chi^2-divergence. In plain English, this measures how "mismatched" your data sources are compared to the real city.

  • If your data is very mismatched, your "Effective Sample Size" is low (it's like having 1,000 bags of flour but only enough for a small cake).
  • If your data is well-matched, your "Effective Sample Size" is high.

The Winning Strategy:
The paper proves that the best way to spend your money is to:

  1. Collect data in a specific mix that maximizes this "Effective Sample Size" (balancing cost and how much the data needs to be "corrected").
  2. Use a "Post-Stratified Estimator." This is the "smart baker." It takes your messy, biased data, looks at how many people from each group you actually got, and mathematically re-weights them to represent the whole city accurately.

What They Proved

The authors didn't just guess this; they did the heavy math to prove it's the best possible way to do it.

  • Lower Bound: They proved that no other method, no matter how clever, can do better than this "Effective Sample Size" limit. You cannot beat the physics of the problem.
  • Upper Bound: They showed that their specific plan (buying data to maximize this size) actually hits that limit. It is "Minimax Optimal," which is a fancy way of saying: "This is the safest, most efficient strategy possible given the worst-case scenario."

Real-World Examples from the Paper

  • Medical Studies: Imagine a study on a new drug. You have clinics in the city (cheap, but mostly young people) and clinics in the countryside (expensive, but mostly elderly). The drug affects both, but you need to know the average effect for the whole country. The paper tells you exactly how many patients to recruit from each clinic to get the most accurate answer for the least money.
  • Political Polling: If you want to know who will win an election, and you have cheap phone surveys (mostly one demographic) and expensive in-person surveys (another demographic), this method tells you how to split your budget to get the truest picture of the voters.

The Bottom Line

Don't try to force your data to look like the population you are studying. Instead, buy as much data as you can afford from the available sources, and then use a smart mathematical tool to "fix" the bias. The paper provides the exact recipe for how to buy that data and the exact tool to fix it, proving that this combination is the absolute best you can do under a budget.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →