← Latest papers
📊 statistics

A Globally Calibrated Bayesian Optimal Phase II Design for Adaptive Enrichment Trials

This paper proposes a pathwise globally calibrated Bayesian optimal phase II (BOP2) design for adaptive enrichment trials that jointly calibrates decision thresholds for all-comer and biomarker-positive subgroups to rigorously control the global type I error rate while enabling efficient treatment development through pre-specified branching rules.

Original authors: Masahiro Kojima, Hisato Sunami, Masaaki Kuriki

Published 2026-07-21
📖 6 min read🧠 Deep dive

Original authors: Masahiro Kojima, Hisato Sunami, Masaaki Kuriki

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery, but you have a limited budget and a ticking clock. In the world of medicine, this detective is a clinical trial, and the mystery is whether a new drug actually works. Usually, doctors test a drug on a huge group of people, hoping to find a cure. But sometimes, the drug is a dud for the general crowd yet a miracle for a specific group of people who carry a special "biological ID card" (called a biomarker). This is where "adaptive enrichment" comes in: it's like a smart detective who starts by questioning everyone, but if the general crowd doesn't give any clues, the detective immediately switches gears to focus only on the people with that special ID card. The goal is to find the cure without wasting time or money on dead ends.

However, there is a tricky trap in this strategy. If you test the general group and the special group separately using two different rulebooks, you might accidentally trick yourself into thinking the drug works when it doesn't. It's like flipping a coin twice; if you keep flipping until you get a "heads," you might convince yourself the coin is magical, even though you just got lucky. Scientists need a way to run this two-step detective story without cheating the odds. This paper introduces a new, mathematically perfect rulebook to make sure the detective stays honest, no matter how the story unfolds.


The Two-Path Detective Story

Imagine a clinical trial as a branching adventure game. You start at the beginning with a massive crowd of players (the "all-comer" population). The game has checkpoints where you check the score. If the score is too low, the game ends for the whole crowd. But here's the twist: if the score is low for the whole crowd, the game doesn't necessarily end. Instead, it might switch to a "secret level" where only players with a specific badge (the "biomarker-positive" subgroup) get to keep playing.

The problem the authors, Masahiro Kojima and his team, are solving is that if you design the rules for the "crowd level" and the "secret level" separately, you might end up with a game that is too easy to win by accident. If you set the bar for the crowd at one height and the bar for the secret level at another, you might accidentally claim victory just because you got lucky on one of the two paths. This is called inflating the "false-positive" rate—thinking you found a cure when you actually didn't.

The "Globally Calibrated" Solution

The paper proposes a new way to set the rules, which they call a "Globally Calibrated Bayesian Optimal Phase II Design." Think of this as designing the entire adventure map at once, rather than drawing the first half and then the second half separately.

Instead of having two separate rulebooks, the authors created one master rulebook that looks at the entire journey. They ask: "What is the chance that we claim a win on either the crowd path or the secret path if the drug is actually useless?" They then adjust the difficulty of the checkpoints (the thresholds) so that this total chance of a false win stays below a strict limit (10% in their example).

Here is how the magic happens:

  1. The Crowd Path: The trial starts with everyone. At specific checkpoints (after 20, 30, and 40 patients), they check if the drug is working. If the drug looks bad, they stop the crowd path.
  2. The Switch: If the crowd path fails, the trial doesn't stop. It immediately switches to the "secret level" with only the biomarker-positive patients.
  3. The Secret Path: This path has its own checkpoints (after 20, 30, and 40 biomarker-positive patients).
  4. The Global Calibrator: The authors used a clever mathematical trick (called "exact finite-state recursive enumeration") to calculate every single possible way the game could play out. They didn't just guess or simulate a few times; they counted every single possibility. This allowed them to tune the rules so that the combined risk of a false alarm stays perfectly under control, no matter how many people have the biomarker.

What They Found

The authors tested their new design against the old way of doing things (where the two paths are calibrated separately). They ran the numbers for different scenarios, changing how common the biomarker was (from 40% to 80% of the population) and how well the drug worked.

The results were clear:

  • The Old Way: When they calibrated the two paths separately, the chance of a false alarm jumped up to between 12.5% and 14.5%. That's like having a 1 in 7 chance of thinking a fake drug is real, which is too risky for medicine.
  • The New Way: With their globally calibrated design, the false alarm rate stayed safely between 8.5% and 9.0%, well under the 10% safety limit.

They also looked at what happens when the drug does work. They found that the new design is just as good at finding real cures as the old one. If the drug works great for the biomarker group, the trial finds it. If the drug works for everyone, the trial finds that too. The only difference is that the new design doesn't cheat the odds.

Why This Matters

This paper doesn't just suggest a new idea; it provides a precise, mathematically proven method to run these complex trials. By treating the "crowd" and the "subgroup" as parts of a single, connected story, the authors ensure that scientists can be confident in their results. They showed that you can be flexible—switching from a broad search to a targeted hunt—without losing your scientific integrity.

In the end, this design is like a masterfully written rulebook for a board game. It allows players to change strategies mid-game based on the clues they find, but it guarantees that the game is fair and that no one can claim a win just because the rules were sloppy. For doctors and patients, this means more reliable answers about which treatments actually work and for whom.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →