← Latest papers
📄 health informatics

Separation-like irregularity and sample size optimism in high-discrimination logistic prediction models

This study demonstrates that closed-form sample size criteria for logistic prediction models, such as the Riley framework, become increasingly optimistic and underestimate the required sample size as target discrimination rises, primarily due to separation-like behavior that destabilizes calibration, thereby necessitating the use of simulation-based stress tests for high-discrimination scenarios.

Original authors: Liu, Z., Liang, Y., Wang, L. S., Yu, J., Liu, J.

Published 2026-01-23
📖 5 min read🧠 Deep dive

Original authors: Liu, Z., Liang, Y., Wang, L. S., Yu, J., Liu, J.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are a chef trying to create a new recipe for a soup that predicts whether a customer will love it or hate it. You have a list of ingredients (predictors) like salt, pepper, and herbs. Your goal is to write down a recipe (a mathematical model) that is so good it can perfectly guess a customer's reaction just by looking at the ingredients.

This paper is about figuring out how many taste tests (data points) you need to write a reliable recipe before you open your restaurant.

The Old Rulebook vs. The New Reality

For a long time, chefs (researchers) used a simple rule of thumb: "If you have 10 ingredients, you need at least 100 taste tests." Later, a more sophisticated rulebook (called the Riley framework, used in a tool called pmsampsize) was created. This rulebook is like a smart calculator that says, "Based on how complex your soup is and how well you think it will taste, here is the exact number of tests you need."

The paper's authors asked a critical question: Is this smart calculator always right?

They discovered that the calculator works great when the soup is "okay" (moderate discrimination). But when the soup is incredibly delicious (high discrimination, meaning the ingredients make the outcome very obvious), the calculator starts to lie. It tells you, "You only need a few tests!" when, in reality, you actually need many more.

The "Perfect Storm" Analogy

Think of it like trying to find a needle in a haystack.

  • Low Discrimination: The needle is buried deep in a huge haystack. It's hard to find. The calculator says, "You need a lot of time to search." This is correct.
  • High Discrimination: The needle is glowing neon and sitting right on top of the hay. It's super easy to spot. The calculator looks at this and says, "Wow, that's easy! You only need 5 seconds to find it."

Here is the problem: Just because the needle is glowing doesn't mean you can stop searching after 5 seconds. If you stop too early, you might miss the exact spot where the needle is, or you might think you found it when you actually just saw a shiny piece of foil.

In the world of statistics, when a model is "too good" at predicting the outcome, the math gets shaky. The calculator assumes the math is smooth and easy, but in reality, the numbers start to go wild (a phenomenon called separation). The model starts predicting "100% chance of love" or "0% chance of love" for almost everyone. This looks great on paper, but it's actually a sign that the recipe is unstable and will fail when you try it on new customers.

What the Authors Did

The authors ran a massive computer simulation (a "virtual kitchen") to test this.

  1. They created thousands of fake soups with different levels of "deliciousness" (discrimination).
  2. They asked the calculator how many taste tests were needed.
  3. They then actually ran the tests to see how many were really needed to get a stable, reliable recipe.

The Result:

  • Moderate Soup: The calculator was right.
  • Super-Delicious Soup: The calculator was way too optimistic. It suggested sample sizes that were 20% to 40% too small. If researchers followed the calculator's advice for these "perfect" soups, their models would be unstable and their risk predictions would be dangerously inaccurate.

The "Glitch" in the System

Why does the calculator fail? The authors found that when the model is too good, the math inside the computer starts to glitch. It's like a GPS that gets confused when you drive too fast; it thinks it knows the route perfectly, but it's actually spinning in circles.

In their simulations, they saw that at the sample sizes the calculator recommended, the computer models often started producing "extreme" results (predicting 99.999% certainty). Even though the computer said "I'm done, I found the answer," the answer was actually a fluke. The model hadn't truly learned the pattern; it had just memorized the noise.

The Takeaway for Chefs (Researchers)

The paper doesn't say "throw away the calculator." It says: Use the calculator as a starting point, but don't trust it blindly when the soup is too perfect.

If you are building a model that you expect to be very accurate (high discrimination), you should:

  1. Run a "Stress Test": Before you collect all your data, run a small simulation to see if the model starts acting crazy (predicting 100% or 0% too often).
  2. Get More Data: If the stress test shows the model is unstable, you need to collect more data than the calculator suggested.
  3. Use a Safety Net: Instead of using a standard recipe, use a "stabilized" recipe (like Ridge regression) that prevents the model from getting too excited and making extreme predictions.

Summary

The paper warns that when a prediction model is too good at telling the future, the standard math used to plan the study breaks down. It tells researchers to be careful not to be fooled by high accuracy; they need more data and better checks to ensure their model is truly reliable and not just a lucky guess.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →