← Latest papers
📊 epidemiology

BUDS: Benchmark Uncertainty Design Selection for Two-Stage Single-Arm Phase II Trials

The paper introduces BUDS, a framework for selecting two-stage single-arm Phase II trial designs that accounts for historical benchmark uncertainty by optimizing efficiency across a plausible range of outcomes, thereby preventing Type I error inflation while offering explicit trade-offs between robustness and sample size efficiency for both binary and time-to-event endpoints.

Original authors: Irlmeier, R., Jin, Z., Ye, F.

Published 2026-07-29
📖 6 min read🧠 Deep dive

Original authors: Irlmeier, R., Jin, Z., Ye, F.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are a detective trying to catch a thief, but you don't know exactly what the thief looks like. You have a sketch, but it's a bit fuzzy. In the world of medical science, specifically in the early stages of testing new drugs, researchers play a similar game. They want to know if a new treatment works better than the "old way" of doing things. To do this, they run what are called "Phase II trials." Think of these as the first serious audition for a new actor. The researchers compare the new drug's performance against a "historical benchmark"—a standard number based on how patients usually do with the current best treatment. It's like saying, "If the new drug cures more than 10% of people, we'll keep going; if not, we stop."

The problem is, that "10%" number is often a guess. It might come from a small study, or the patients might be slightly different, or medicine might have changed a little since the old study was done. If the researchers pick the wrong number for their benchmark, they might accidentally think a useless drug is a hero, or they might throw away a good one too soon. This paper tackles the tricky math of how to design these drug auditions when you aren't 100% sure what the "standard" performance actually is. It's about building a safety net so that even if your guess is a little off, you don't make a costly mistake.


The "BUDS" Solution: Designing for the Unknown

This paper introduces a new framework called BUDS (Benchmark Uncertainty Design Selection). Think of BUDS as a smart, cautious architect designing a bridge. Usually, when you build a bridge, you calculate the load based on a single, specific weight—say, a car that weighs exactly 2,000 pounds. You design the bridge to hold that car perfectly. But what if, in reality, the cars crossing the bridge could weigh anywhere between 1,800 and 2,200 pounds? If you only designed for 2,000, a heavy truck might snap your bridge.

The authors argue that traditional drug trial designs are like that single-weight bridge. They pick one specific "benchmark" (like a 10% cure rate) and design the whole study around it. The paper shows that if the real-world benchmark is actually higher (say, 15%), these traditional designs can fail spectacularly. They might declare a drug a success when it's actually a dud. In their simulations, the authors found that when the true benchmark was just a bit higher than planned, the error rate for a popular design called "Simon Optimal" skyrocketed to 0.407 (that's a 40.7% chance of a false alarm!) instead of the intended 10%.

To fix this, BUDS suggests stopping the "single guess" game. Instead of picking one number, researchers should define a plausible range of what the benchmark could be (e.g., "The cure rate is likely between 5% and 15%"). BUDS then designs the trial to work safely across that entire range.

How BUDS Works: The "Worst-Case" and "Average" Strategies

The paper proposes two main ways to pick the best trial design from the many options that fit this new, wider safety net:

  1. BUDS-Worst-Regret: This is the "paranoid" strategy. It asks, "What is the worst-case scenario within our range, and how much extra work (patients) would we need to handle that specific worst case?" It then picks the design that minimizes the maximum disappointment or "regret" if the worst case turns out to be true. It's like packing a suitcase for a trip where you might need to carry a heavy coat, a light jacket, or nothing at all. You choose the outfit that keeps you comfortable even if you end up in the coldest weather, even if it means you're slightly overdressed for a sunny day.
  2. BUDS-Avg-EN: This is the "average" strategy. It looks at the whole range of possibilities and picks the design that uses the fewest patients on average. It's like packing for a trip where you expect a mix of weather and just want to be efficient overall, even if you might get a little chilly on one specific day.

What the Simulations Showed

The authors ran thousands of computer simulations to see how these new designs compare to the old ones. Here is what they found:

  • Safety First: The most important finding is that BUDS designs keep the "false alarm" rate (Type I error) under control no matter where the true benchmark falls within the range. Whether the benchmark is at the low end or the high end of the guess, the trial won't accidentally declare a bad drug a winner. In contrast, the old single-benchmark designs saw their error rates explode when the real benchmark was higher than expected.
  • The Cost of Safety: There is a trade-off. To be safe across a wide range of possibilities, BUDS designs sometimes require more patients than the old designs. For example, in a scenario where the benchmark was uncertain, a traditional design might need a maximum of 37 patients. The BUDS design, to be safe against the worst-case scenario, might need 54 or even 72 patients. The paper notes that this extra enrollment is the price of not making a mistake. It's better to test a few extra people than to waste millions on a Phase III trial for a drug that doesn't work.
  • Real-World Examples: The authors tested their ideas on real-world scenarios, like a trial for glioblastoma (a type of brain cancer). They looked at a trial that used a benchmark of 0.10 (10%). When they applied BUDS with a realistic range of 0.05 to 0.13, the old design would have had a false alarm rate of over 0.25 (25%). The BUDS design kept the error rate below 0.093 (9.3%), but it required a maximum sample size of 89 patients instead of the original 50.

The Bottom Line

The paper concludes that BUDS offers a transparent way to handle the uncertainty of medical history. It doesn't magically make the uncertainty disappear; instead, it forces researchers to admit, "We aren't 100% sure, so let's plan for a range." By doing so, it makes the trade-off between efficiency (using fewer patients) and robustness (avoiding false positives) explicit. The authors have even built a free computer tool (an R package and a Shiny app) so other scientists can use this method to design their own trials.

In short, BUDS suggests that in the high-stakes game of drug testing, it's smarter to design a bridge that can handle a range of weights than to gamble on a single, perfect guess. The paper doesn't claim this is the final word on all medical trials, but it provides a solid, mathematically sound way to stop guessing and start planning for the unknown.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →