← Latest papers
📊 statistics

Bayesian sample size calculations for external validation studies of risk prediction models

This paper proposes a general Bayesian framework for determining sample sizes in external validation studies of risk prediction models that explicitly accounts for uncertainty in model performance and incorporates multi-criteria considerations, including expected precision, assurance probabilities, and Value of Information for net benefit, often resulting in more efficient sample size requirements compared to conventional methods.

Original authors: Mohsen Sadatsafavi, Paul Gustafson, Solmaz Setayeshgar, Laure Wynants, Richard D Riley

Published 2026-02-13
📖 5 min read🧠 Deep dive

Original authors: Mohsen Sadatsafavi, Paul Gustafson, Solmaz Setayeshgar, Laure Wynants, Richard D Riley

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a chef who has developed a new, revolutionary soup recipe. You've tested it in your home kitchen (the development study) and it tasted amazing. Now, you want to serve this soup to a whole new city of people (the target population). Before you open a restaurant there, you need to run a "taste test" to make sure the soup still works well with different ingredients and different palates. This is what scientists call external validation.

The big question is: How many people do you need to invite to this taste test?

If you invite too few, you might get a fluke result (maybe the soup just happened to taste good that day). If you invite too many, you waste time, money, and ingredients.

The Old Way: Guessing the Exact Flavor

Traditionally, scientists used a method (like the one by Riley et al.) that asked you to guess a single, fixed number for how good the soup is.

  • Example: "I bet the soup will have a 90% success rate."
  • Then, they calculated: "To prove it's 90% good with a margin of error of 5%, we need 1,000 people."

The Problem: This approach ignores reality. You don't actually know the soup is exactly 90% good. Maybe it's 85%, maybe 95%. The old method treats your guess as absolute truth, which is risky. It's like planning a road trip assuming you will drive exactly 60 mph the whole time, ignoring traffic, wind, and fatigue.

The New Way: The Bayesian "Weather Forecast"

This paper proposes a smarter, Bayesian approach. Instead of guessing one fixed number, you admit you are uncertain and create a range of possibilities (a probability distribution).

Think of it like a weather forecast.

  • Old Way: "It will rain at 2:00 PM." (Too specific, often wrong).
  • New Way: "There is a 70% chance of rain, a 20% chance of drizzle, and a 10% chance of sun." (Honest about uncertainty).

In this new framework, you say: "Based on previous tests, I'm pretty sure the soup is good, but it could be anywhere between 'okay' and 'amazing'." You then run thousands of computer simulations to see what happens with different sample sizes.

Three New Rules for the Taste Test

The paper suggests three different ways to decide how many people to invite, depending on what you care about:

1. The "Average" Rule (Expected Precision)

  • The Analogy: You want the average result of your taste test to be very precise.
  • How it works: You ask, "If I invite 500 people, what is the average width of my error bar?" If the average error is small enough, you're happy.
  • Why it's better: It uses all the data you have, not just a single guess.

2. The "Safety Net" Rule (Assurance)

  • The Analogy: You are a risk-averse investor. You don't just want the average result to be good; you want to be 90% sure that you won't get a bad result.
  • How it works: You ask, "If I invite 500 people, is there a 90% chance that my error bar will be small enough?"
  • Why it's better: Sometimes the "average" is good, but there's a small chance of a disaster. This rule protects you from that disaster. It's like buying extra insurance.

3. The "Is It Worth It?" Rule (Value of Information)

  • The Analogy: This is the most creative part. Imagine you are deciding whether to open a restaurant chain.
    • Strategy A: Don't open the restaurant (Safe, but no profit).
    • Strategy B: Open the restaurant based on the soup (Risky, but high profit if good).
    • Strategy C: Open the restaurant based on a different soup.
  • The Question: "If I invite 500 people to taste the soup, will that information actually help me make a better business decision?"
  • How it works: If the taste test results are likely to change your mind (e.g., "Oh wow, the soup is actually terrible, let's not open!"), then the test is valuable. If the test results won't change your mind (you were going to open it anyway, or you were going to cancel it anyway), then the test is waste of money, no matter how precise it is.
  • The Surprise: In the paper's example (predicting COVID-19 patient decline), they found that to get a "statistically perfect" measurement of calibration, they needed 1,181 patients. But, if they just wanted to know if the model was useful for making medical decisions, they only needed 522 patients. The extra 600 patients didn't add enough value to justify the cost.

The Real-World Example: The COVID-19 Soup

The authors tested this on a real model used to predict if hospitalized COVID-19 patients would get worse.

  • The Old Method said: "You need 1,056 patients to be statistically precise."
  • The New Method said:
    • If you want to be 90% sure you get a precise measurement: "You need 1,181 patients."
    • If you ask, "Will testing more people actually change our medical decisions?" (Value of Information): "No. 522 patients is enough."

The Conclusion: By using the "Value of Information" approach, the researchers realized they could save hundreds of patients' time and resources without losing any real-world benefit.

Why This Matters

This paper is a call to stop treating scientific models like perfect, unchangeable facts.

  1. Acknowledge Uncertainty: Admit that we don't know the exact truth yet.
  2. Be Flexible: Choose sample sizes based on how much risk you are willing to take (Assurance).
  3. Be Practical: Don't just collect data for the sake of "precision." Collect data only if it helps you make a better decision (Value of Information).

In short, this new method is like switching from a rigid, blindfolded walk to a guided hike where you check the map, check the weather, and ask, "Is this detour actually going to get us to the destination faster?"

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →