← Latest papers
📊 statistics

Evaluating the impact of outcome delay on the efficiency of sample size re-estimation

This paper evaluates how outcome delays in clinical trials with sample size re-estimation reduce efficiency by inflating the average sample size and power, particularly when the re-estimated sample size is smaller than originally planned, leading to costly overpowered trials.

Original authors: Aritra Mukherjee, Michael J Grayling, James J M S Wason

Published 2026-05-14
📖 5 min read🧠 Deep dive

Original authors: Aritra Mukherjee, Michael J Grayling, James J M S Wason

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are organizing a massive party to test a new recipe. You need to invite enough people to be sure the recipe is a hit, but you don't want to invite too many, or you'll waste money on food that nobody eats.

In the world of clinical trials (testing new medicines), scientists face this exact problem. They have to guess how many patients to enroll to prove a drug works. This guess depends on a "nuisance parameter"—a fancy way of saying "a variable we aren't sure about," like how much the results might vary from person to person.

The Solution: The "Internal Pilot" (The Taste Test)
To fix their guess, researchers use a strategy called Sample Size Re-estimation (SSR). Think of this as a "taste test" halfway through the party.

  1. They invite a small group of people (the first stage).
  2. They check the data to see how variable the results are.
  3. Based on that, they adjust the total number of guests they need. If the results are very consistent, they might need fewer people. If they are all over the place, they might need more.

The Problem: The "Pipeline" (The Slow Cook)
Here is where things get tricky. In many medical trials, the "result" (did the patient get better?) doesn't happen immediately. It might take months to see if the drug worked. This is called Outcome Delay.

Imagine you are the party host. You've invited 60 people for the taste test. You are waiting for the results of their feedback before you decide how many more people to invite.

  • The Ideal Scenario: You stop inviting new people while you wait for the feedback.
  • The Real Scenario: You keep inviting people because the invitation process is automatic. While you are waiting for the 60th person's feedback (which takes 8 months), you keep inviting new guests.

By the time the feedback arrives, you might have invited 100 people total. But the taste test told you that you only needed 80 people to get a good answer. You now have 20 "pipeline" guests—people who are at the party but whose feedback won't count toward your decision. You've wasted resources on an over-powered trial (too many people, too much cost).

What This Paper Found
The authors of this paper ran thousands of computer simulations to see how bad this "waiting game" actually is. They used three main ways to measure the damage:

  1. The "Delay Impact" (How often do we overshoot?):

    • They found that the longer you wait for the results, the more likely you are to accidentally invite too many people.
    • The Big Surprise: This problem is worst when the drug works better than expected (or the data is less variable than you thought). In these cases, the "taste test" tells you to cut the guest list down. But because you kept inviting people while waiting, you end up with a huge crowd anyway.
    • Conversely, if the data is messy and you need more people, the delay isn't as bad. The extra people you invited while waiting actually help you reach the higher number you needed.
  2. The "RMSE" (How far off are we?):

    • This measures how far the final number of guests is from the "perfect" number.
    • The study shows that as the delay gets longer, the final number of guests drifts further away from the ideal. The trial becomes less efficient.
  3. The "Cost" (The Penalty Score):

    • The authors created a special score that punishes "under-powered" trials (not enough people to prove the drug works) more severely than "over-powered" ones.
    • They found that while delays usually lead to over-powering (wasting money), the specific "cost" of the delay depends heavily on the initial guess. If you guessed the variance was high (needed many people) but it turned out to be low (needed few), the delay causes a massive spike in cost because you recruited way more than necessary.

The Recipe for Success
The paper concludes with a specific recommendation on how to handle this "waiting period":

  • Don't wait too long to check: The timing of your "taste test" matters.
  • The Sweet Spot: The authors suggest that for continuous data (like blood pressure readings), checking the data after 35 people per group is the best balance.
    • If you check too early (fewer people), your guess might be wild and inaccurate.
    • If you check too late (many people), you risk inviting a huge "pipeline" of unnecessary guests while you wait for the results.

In a Nutshell
Delaying the results of a clinical trial while you keep recruiting patients is like cooking a stew and keeping the pot full of water while you wait for the soup to boil. If you need less water than you thought, you've wasted a lot of it. This paper shows that this waste is most dangerous when the trial turns out to be "easier" than expected. To avoid this, researchers should do their "taste test" (re-estimation) after about 35 patients per group, rather than waiting until they have recruited a huge crowd.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →