← Latest papers
📊 statistics

Sample Size Determination Under Selection Bias: Robust Tolerance Limits for Prevalent Cohort Data

This paper derives and validates modified distribution-free tolerance limit formulae that accommodate various biased sampling schemes, such as weight bias and censoring, to address the limitations of traditional methods when unbiased representative samples are unavailable, demonstrating their application with dementia data from the Canadian Study of Health and Aging.

Original authors: James H. McVittie, Martin Lysy, Masoud Asgharian

Published 2026-05-20
📖 4 min read☕ Coffee break read

Original authors: James H. McVittie, Martin Lysy, Masoud Asgharian

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a baker trying to figure out how many cookies to bake to ensure that, if you grab two specific cookies from the tray (say, the 3rd smallest and the 10th largest), the space between them contains at least 80% of all the cookie sizes you could have made.

In the world of statistics, this is called finding a tolerance limit. For decades, statisticians have had a famous, simple recipe (the Scheffé-Tukey formula) to answer this question. The recipe works perfectly if your tray of cookies is a fair, random sample of the whole batch. If you pick cookies blindly, the math holds up.

The Problem: The "Biased" Tray
However, in real life—especially in medical studies like tracking dementia patients—you often can't get a fair, random tray. You might only get cookies that are "longer" or "heavier" because of how you collected them.

  • The Analogy: Imagine you are studying how long people live with a disease. If you only start looking at people after they have already been sick for a while (a "prevalent cohort"), you are more likely to find the people who survive longer. The short-lived patients have already passed away and are invisible to you.
  • The Result: Your sample is biased. It's like a tray where only the biggest cookies were allowed to stay. If you use the old, simple recipe on this biased tray, you will calculate that you need very few cookies to get your 80% coverage. But in reality, you will be way off. You'll think you have enough data, but you won't.

The Solution: New Recipes for a Biased World
The authors of this paper, James McVittie, Martin Lysy, and Masoud Asgharian, realized that the old recipe fails when the data is biased. They created two new "recipes" (methods) to fix the math so it works even when your sample is skewed.

  1. The "Inequality" Method (The Over-Preparer):
    This method tries to be safe by using a series of logical "if-then" rules. It's like a baker who, worried about the biased tray, decides to bake way more cookies than necessary just to be absolutely sure.

    • The Catch: It works, but it's wasteful. It tells you to collect a massive amount of data (a huge sample size) to be safe, often overestimating what you actually need.
  2. The "FFT" Method (The Precision Chef):
    This is the paper's star player. "FFT" stands for Fast Fourier Transform, which is a fancy math tool that acts like a high-speed blender for probability curves.

    • How it works: Instead of guessing or using loose rules, this method takes the shape of your biased data and mathematically "un-blends" it to see what the original, fair distribution looked like. It then calculates the exact number of cookies (sample size) needed.
    • The Result: It is incredibly accurate. It tells you the exact number of people you need to study to get your 80% coverage, no more and no less.

The Proof: The Simulation and the Real Test
The authors tested these new recipes in two ways:

  • The Simulation: They created fake computer data where they knew the "true" answer. When they used the old recipe on biased data, it failed miserably (predicting they needed too few people). The "Inequality" method was too cautious (predicting too many). The FFT method hit the bullseye every time.
  • The Real World Test: They applied this to real data from the Canadian Study of Health and Aging (CSHA), which tracks people with dementia. Because this study only catches people who have already been diagnosed (a biased sample), the old math would have given the wrong answer. Using their new FFT method, they calculated exactly how many more participants would be needed to be statistically confident.

The Bottom Line
If you are doing a study where your data is naturally "biased" (like looking at people who have already survived a disease for a while), do not use the old, simple math formula. It will lie to you.

Instead, use the new FFT method proposed in this paper. It's like upgrading from a rough guess to a laser-guided calculator. It ensures you don't waste money and time collecting too much data, nor do you risk failing your study by collecting too little. It's the most efficient way to handle messy, real-world data.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →