← Latest papers
📊 statistics

Bayesian sample size determination using robust commensurate priors with interpretable discrepancy weights

This paper addresses the nonmonotonicity and interpretability issues in Bayesian sample size determination using commensurate priors by proposing a linearization technique that yields an analytical formula and interpretable weights representing the probability of historical data relevance, thereby facilitating expert elicitation for trial design.

Original authors: Lou E. Whitehead, James M. S. Wason, Oliver Sailer, Haiyan Zheng

Published 2026-06-25
📖 5 min read🧠 Deep dive

Original authors: Lou E. Whitehead, James M. S. Wason, Oliver Sailer, Haiyan Zheng

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: Cooking with Old Recipes

Imagine you are a chef planning a new, high-stakes cooking competition. You need to decide how many judges (participants) to invite. If you invite too few, you might not get a clear winner; if you invite too many, you waste money and time.

Usually, chefs rely on new data to make this decision. But what if you have old cookbooks (historical studies) from similar competitions? It would be smart to use those old recipes to help you decide how many judges you need. This is the core idea of the paper: using past data to design a new medical trial more efficiently.

However, there is a catch. You can't just copy-paste the old recipes. You have to ask: "How similar is this old recipe to what I'm cooking today?"

The Problem: The "Volume Knob" is Broken

In the past, statisticians used a "volume knob" (called a discrepancy weight) to control how much of the old data to listen to.

  • Turn the knob to 0: Listen to the old data 100%.
  • Turn the knob to 1: Ignore the old data completely.

The authors discovered that in previous methods, this volume knob was broken.

  1. It wasn't linear: If you turned the knob just a tiny bit (from 0 to 0.1), the volume dropped drastically. But if you turned it from 0.5 to 0.6, the volume barely changed. It was like a light switch that was too sensitive at the start and useless at the end.
  2. It wasn't predictable: Sometimes, turning the knob up (to ignore data) actually made the system listen more to the data. This is called non-monotonicity. It's like trying to turn down the heat on a stove, but the flame suddenly gets bigger.

Because of this broken knob, it was impossible for experts to say, "I think this old study is 50% relevant," because the math wouldn't actually give them 50% relevance. It made communication between statisticians and doctors very difficult.

The Solution: A New, Smooth Knob

The authors propose a two-step fix to make the knob work properly.

Step 1: A Better Way to Mix the Ingredients

They changed the mathematical recipe for how the old data is mixed together. Instead of the messy, unpredictable mixing method used before, they used a cleaner method (based on the work of Winkler, 1981).

  • The Result: Now, the relationship between the knob and the volume is smooth and predictable. Turning the knob up always lowers the volume. No more surprises.

Step 2: The "Magic Translator"

Even with the new mixing method, the relationship between the knob and the final result (the number of judges needed) is still a bit curved.

  • The Fix: They created a "Magic Translator" (a mathematical transformation).
  • How it works: You tell the translator, "I want to use 50% of the old data." The translator takes that 50%, does a quick calculation, and tells the computer, "Okay, set the internal knob to this specific tiny number."
  • The Benefit: Now, when an expert says, "I'm 50% skeptical of this old data," the system actually uses exactly 50% of the information. The relationship becomes linear (a straight line).

The Real-World Test: Alzheimer's Exercise Study

To prove this works, the authors applied their method to a hypothetical study about Alzheimer's disease.

  • The Goal: Test if exercise helps improve memory.
  • The Data: They looked at 7 past studies. Some were very similar to the new plan; others were quite different.
  • The Expert: A doctor looked at the 7 studies and said, "Study #5 is very relevant (low weight), but Study #6 is not relevant at all (high weight)."
  • The Outcome:
    • Without their new method, the math would have misunderstood the doctor's weights, leading to a trial with 332 participants (too big).
    • With their new "Magic Translator," the math correctly interpreted the weights and calculated a trial size of 176 participants.
    • This saved resources while still ensuring the trial was big enough to be scientifically valid.

Why This Matters

The paper doesn't claim to cure diseases or invent new drugs. Instead, it fixes the toolkit used to design the experiments that test those drugs.

By making the "volume knob" for historical data intuitive and predictable, they hope that:

  1. Doctors and Statisticians can talk better: Doctors can give honest opinions on how relevant old data is without worrying the math will break.
  2. Trials are more efficient: We can stop wasting money on trials that are too big, or risking failure on trials that are too small.

In short, they built a better ruler for measuring how much we should trust the past when planning the future.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →