← Latest papers
📊 statistics

Incorporating Missing Data Considerations into Sample Size Calculations for Developing Clinical Prediction Models

This study demonstrates that missing predictor data significantly increases the sample size requirements for developing stable and well-calibrated clinical prediction models, and proposes a practical adaptation of posterior-based sample size calculations to directly incorporate missing data assumptions and imputation strategies.

Original authors: Glen P. Martin, Sian Bladon, Rebecca Whittle, Molly Wells, Gary S. Collins, Richard D. Riley

Published 2026-05-11
📖 6 min read🧠 Deep dive

Original authors: Glen P. Martin, Sian Bladon, Rebecca Whittle, Molly Wells, Gary S. Collins, Richard D. Riley

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: Why We Need More Data When Things Go Missing

Imagine you are a chef trying to create a new, perfect soup recipe (a Clinical Prediction Model). To make sure the soup tastes right for everyone, you need to taste-test it on a large group of people.

Usually, statisticians have a "rule of thumb" for how many people you need to taste-test to get a reliable recipe. They assume that every single person you ask will have all the ingredients ready to go (e.g., they all have salt, pepper, and carrots).

The Problem: In the real world, people often forget to bring an ingredient. Maybe someone has no carrots, or their salt is missing. This is missing data.

This paper argues that if you expect people to be missing ingredients, your old "rule of thumb" for how many people to taste-test is wrong. It's too low. If you stick to the old number, your soup might taste great in your small test kitchen, but it will be a disaster when you try to serve it to the whole world.

The Core Discovery: Missing Ingredients Need a Bigger Pot

The researchers ran a massive simulation (a computer experiment) to see what happens when you try to cook with missing ingredients.

  1. The "Perfect" Scenario: If everyone brings all their ingredients, the standard recipe for the number of testers works fine.
  2. The "Missing" Scenario: When 20%, 40%, or even 60% of the testers are missing ingredients, the recipe starts to fail. The soup becomes "over-fitted."
    • Analogy: Think of "over-fitting" like memorizing a specific test. If you only study the exact questions on one practice test, you might get 100% on that test. But if you walk into the real exam and the questions are slightly different (or some are missing), you fail. The model has learned the "quirks" of the small, messy dataset rather than the true pattern.

The Finding: To get the same reliable soup recipe when ingredients are missing, you often need twice as many testers as the standard rules suggest.

The Solution: A New "Simulation Kitchen"

The paper proposes a new way to figure out how many people you need. Instead of just doing a quick math calculation, they suggest running a simulation (a practice run) before you even start your real study.

Here is how their new method works, step-by-step:

  1. Make a Fake World: First, you create a giant, perfect computer world of 500,000 people where you know exactly what the "true" soup recipe should be.
  2. Break the World: Then, you intentionally "break" this world by making some people forget their ingredients (simulating missing data).
  3. The Practice Run: You try to cook your soup using different strategies:
    • Strategy A: Throw away anyone missing an ingredient (Complete Case Analysis).
    • Strategy B: Guess the missing ingredient based on what they did bring (Imputation).
  4. Check the Taste: You taste the soup. Does it match the "true" recipe?
  5. Adjust the Pot Size: If the soup tastes bad, you increase the number of testers in your simulation and try again. You keep doing this until you find the exact number of testers needed to guarantee the soup will taste good, even with missing ingredients.

Key Concepts Explained Simply

  • Overfitting: Imagine a student who memorizes the answers to a specific practice quiz. They get an A. But when the real test comes, and one question is missing, they panic and fail. The model is "overfit" because it learned the noise (the missing data quirks) instead of the signal (the real pattern).
  • Calibration Slope: This is a score that tells you how "shrink-wrapped" your model is. A score of 1.0 is perfect. A score below 0.9 means your model is too confident and likely to be wrong. The paper found that with missing data, scores often dropped below 0.9 unless you doubled your sample size.
  • Imputation: This is the act of "guessing" the missing values. The paper tested different ways of guessing (like using a random forest or a simple average) and found that while guessing helps, it doesn't fix the problem entirely—you still need more data.
  • EVPI (Expected Value of Perfect Information): This is a fancy way of asking, "How much are we losing by not having perfect data?" It calculates the cost of uncertainty. The paper shows that missing data increases this "cost," and adding more people to your study lowers it.

What the Paper Actually Says (and Doesn't Say)

  • What they found: Missing data makes your model less stable and less accurate. To fix this, you need to increase your sample size, sometimes by a factor of 2.
  • What they recommend: Don't just use the old, simple math formulas. Use their new simulation method to plan your study. This involves making assumptions about how data goes missing (e.g., is it random, or does it happen because of other factors?) and testing those assumptions in a computer simulation.
  • What they didn't do: They did not test this on a specific real-world hospital right now. They used computer simulations and two hypothetical examples (one where they had old data to help them guess, and one where they had to guess blindly).
  • The "Rule of Thumb" Alternative: If you can't do the complex simulation, they suggest a rough fix: take your standard sample size number and divide it by (1 - percentage of missing data). For example, if you expect 50% missing data, you need double the sample size. However, they admit this is just a rough guess and the simulation is better.

The Bottom Line

If you are planning a study to build a medical prediction model, and you know that some data will be missing (which is almost always the case), do not trust the standard sample size calculators. They will give you a number that is too small.

Instead, use this new "simulation kitchen" approach. It allows you to test different scenarios and find the real number of people you need to recruit to ensure your model is stable, accurate, and ready for the real world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →