← Latest papers
📊 statistics

Two-Stage Design with Sample Size Re-estimation Using gsDesign

This paper demonstrates the derivation of two-stage group sequential and conditional power designs using the `gsDesign` R package, arguing that while conditional power designs offer some flexibility, standard group sequential designs are generally preferable due to the risk of inefficiency and potential unblinding of interim treatment effects.

Original authors: Keaven M. Anderson, Alison Pedley

Published 2026-08-05
📖 6 min read🧠 Deep dive

Original authors: Keaven M. Anderson, Alison Pedley

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery, but you don't know how many clues you'll need to catch the culprit. In the world of medical science, this is the daily challenge of designing a clinical trial. Researchers are testing a new medicine to see if it works better than the old one, but they can't be 100% sure how powerful the new drug is before they start. If they guess too low, they might waste time and money testing a drug that's actually a dud. If they guess too high, they might enroll too many patients, exposing them to unnecessary risks and draining the budget. To solve this, scientists use a strategy called "Group Sequential Design." Think of this like a hiking trip with a map that has two checkpoints. You start with a plan, but at the first checkpoint (an interim analysis), you check your progress. If the trail looks amazing, you might stop early because you've already found the treasure. If the trail looks dead, you might turn back to save your energy. But what if the trail is just "okay"? That's where things get tricky.

This paper, written by Keaven M. Anderson and Alison Pedley from Merck & Co., tackles a specific problem with these "checkpoint" trials: what happens when the data is messy, or when the initial guess about the drug's power was wrong? The authors explore a clever workaround called "Conditional Power." Imagine you are halfway through your hike, and you realize the mountain is taller than you thought. Instead of just giving up or blindly marching on, you recalculate: "If I keep going at this pace, what are my chances of reaching the summit?" If the math says your chances are low, you might decide to bring in more supplies (enroll more patients) to boost your odds. The paper uses a computer tool called gsDesign to simulate these scenarios, comparing the traditional "checkpoint" method against this new "recalculate-and-adapt" method. The goal is to see if changing the plan mid-trip actually saves money or if it just makes the journey more complicated and expensive.

The authors set up a virtual experiment to test these ideas. They imagined a trial for a new treatment where the expected benefit was a specific number (0.33), but they worried the real benefit might be smaller (0.27). They started with a standard "Group Sequential" design, which is like a rigid hiking plan with two stops. If the drug looked great at the first stop, they'd stop early. If it looked terrible, they'd stop for futility. If it was just "meh," they'd keep going to the end. However, they noticed a flaw: if the drug was actually only moderately effective, the original plan might not have enough power to prove it works, leading to a failed trial.

To fix this, they tested three different "Conditional Power" strategies. These are like having a smart GPS that recalculates the route at the halfway point based on how well you're actually doing.

  1. The "Observed Effect" Strategy: The GPS looks at your current speed and says, "You're going slower than planned, so let's add more hikers to the team to make sure we reach the top."
  2. The "Planned Effect" Strategy: The GPS ignores your current slow speed and says, "Stick to the original plan; assume you'll speed up later, but just in case, let's add a few more hikers."
  3. The "Single Adapted N" Strategy: The GPS simplifies things by saying, "We only have two options: either we stop now, or we switch to one specific, larger team size."

The researchers ran these simulations to see which approach was the most efficient. They measured two things: the "Power" (the chance of successfully proving the drug works) and the "Expected Sample Size" (the average number of patients needed, which represents the cost).

Here is the twist, and it's a big one: The paper finds that while these "recalculate" strategies sound like a brilliant way to save money, they don't actually work as well as hoped. When the drug's effect was the smaller, harder-to-detect size (0.27), the conditional power designs managed to boost the success rate, but they did so by requiring a significantly larger number of patients on average. In fact, the "Expected Sample Size" for these adaptive designs was often just as high, or even higher, than simply planning for the smaller effect size from the very beginning.

The authors suggest that the "Conditional Power" approach is a bit of a trap. It creates a false sense of security. You think you are being smart by adapting, but you end up paying a high price in terms of the number of patients you need to enroll. The paper explicitly argues against the idea that these designs offer a "smaller up-front planned sample size" as a major advantage. In their simulations, the traditional Group Sequential design, or even a design that simply planned for the worst-case scenario from the start, often came out ahead in terms of efficiency.

Furthermore, the paper points out a hidden risk. By looking at the data to decide whether to add more patients, you might accidentally reveal something about how well the drug is working before the trial is officially over. This could bias the results or make the trial harder to interpret. The authors conclude that while the math behind "Conditional Power" is elegant and the software tools (like gsDesign) make it easy to calculate, the practical benefits are slim. In many cases, the "Group Sequential" design remains the better choice because it is simpler, less likely to be inefficient, and avoids the pitfalls of mid-trial adjustments that can backfire.

So, what's the takeaway for our detective? If you are planning a medical trial, don't get too fancy with mid-trip recalculation unless you are absolutely sure it's necessary. The paper suggests that trying to "rescue" a trial by adding more patients halfway through often costs more than just planning for the worst-case scenario from day one. The most efficient path isn't always the one that looks like it's adapting the most; sometimes, a solid, well-thought-out plan from the start is the best way to reach the summit without burning out your resources. The authors didn't prove that conditional power is useless in every single situation, but they showed that for the scenarios they tested, it was often an inefficient detour compared to sticking with a robust, pre-planned strategy.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →