← Latest papers
📊 statistics

Balancing Evidentiary Value and Sample Size of Adaptive Designs with Application to Animal Experiments

This paper proposes the Experimental Unit Information Index (EUII), a novel metric that balances statistical error rates and sample size to quantify the evidentiary value of individual experimental units, demonstrating its utility in optimizing adaptive animal experiments to reduce sample sizes while maintaining reliable inferences.

Original authors: Leonhard Held, Fadoua Balabdaoui, Saverio Fontana, Samuel Pawel

Published 2026-05-08
📖 5 min read🧠 Deep dive

Original authors: Leonhard Held, Fadoua Balabdaoui, Saverio Fontana, Samuel Pawel

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a scientist running experiments on animals. Your goal is to find out if a new treatment works. But you also have a moral and practical duty: you want to use as few animals as possible (the "Reduce" principle) while still being absolutely sure your results are real and not just a lucky accident.

Usually, scientists play a tricky balancing act. If they use too few animals, they might miss a real effect. If they use too many, they are wasting resources. Traditionally, they compare different experiment designs by looking at two numbers: how often they make a false alarm (Type I error) and how often they successfully find a real effect (Power).

But this paper introduces a new way to look at the problem. It asks: "How much useful information does one single animal actually give us?"

Here is the breakdown of their idea using simple analogies:

1. The Problem with the Old Way

Imagine you are trying to guess if a coin is fair or biased.

  • Method A: You flip the coin 100 times. You get a clear answer.
  • Method B: You flip the coin 10 times, but you cheat a little bit by peeking at the results early and stopping if you see a pattern. You get an answer faster, but your "cheating" makes your error rate slightly higher.

Traditionally, comparing Method A and Method B is hard because they have different rules for "cheating" (error rates) and different costs (number of flips). You can't easily say which one is "better" just by looking at the raw numbers.

2. The New Tool: The "Animal Information Score" (EUII)

The authors created a new score called the Experimental Unit Information Index (EUII). Think of this as a "Value per Animal" score.

To understand this, they borrowed a concept from medical diagnostics called the Diagnostic Odds Ratio.

  • Imagine a medical test for a disease. A "good" test is one where a positive result strongly suggests you have the disease, and a negative result strongly suggests you are healthy.
  • The authors realized that a statistical experiment is just like a medical test. The "disease" is the hypothesis that the treatment works.
  • They calculated a score that tells you: "If I add one more animal to my experiment, how much does it improve my ability to tell the truth from a lie?"

If the score is high, every animal you use is working very hard to give you a clear answer. If the score is low, you are wasting animals on data that doesn't help much.

3. The "Early Stop" Trick (Adaptive Designs)

The paper focuses heavily on Adaptive Designs. This is like a detective who doesn't stick to a rigid plan.

  • The Plan: "I will interview 100 witnesses."
  • The Adaptive Twist: "I will interview witnesses one by one. If I find a smoking gun at witness #20, I stop immediately (Efficacy Stop). If I realize the story makes no sense at witness #30, I stop immediately because it's a waste of time (Futility Stop)."

This saves animals (or witnesses) because you don't interview the rest if you already know the answer. However, stopping early changes the math. It makes the "Value per Animal" calculation tricky because the number of animals used isn't fixed anymore.

4. What They Found

The authors took this new "Value per Animal" score and applied it to thousands of real past animal experiments (about 2,700 of them) to see what would have happened if they had used these "Early Stop" rules.

  • The Big Win: They found that using "Early Stop" rules (especially a specific type called the Pocock design) could save a massive number of animals. In their simulation, they could have saved over 13,000 animals across those studies by stopping early when the results were obvious.
  • The Trade-off: Sometimes, stopping early means you reject the "null hypothesis" (saying the treatment works) slightly less often than a rigid design. But the authors argue this is okay because the quality of the information per animal remaining is higher.
  • The Winner: Among the different "Early Stop" strategies, the Pocock method (which uses a consistent, slightly looser threshold for stopping early) generally gave the highest "Value per Animal" score. It was the most efficient at squeezing information out of every single animal used.

5. The "N-Hacking" Controversy

The paper also looked at a method called "N-hacking" (or "Constrained Sample Augmentation"), proposed by another researcher named Reinagel.

  • The Idea: If the results look "almost significant," add a few more animals to see if it crosses the line. If it looks hopeless, stop.
  • The Catch: This method technically increases the chance of a false alarm (Type I error).
  • The Verdict: The authors found that even with this higher error rate, Reinagel's method was very efficient at saving animals and had a high "Value per Animal" score. However, they noted that the Pocock method (which controls errors strictly) was still a very strong contender and often performed just as well or better depending on the situation.

The Bottom Line

This paper doesn't just say "use fewer animals." It provides a mathematical ruler to measure exactly how much information each animal contributes.

By using this ruler, they showed that stopping experiments early (when the answer is clear or when it's hopeless) is a smart way to follow the "Reduce" principle. It allows scientists to get reliable answers while using fewer animals, without having to sacrifice the reliability of the science. The Pocock design emerged as a top recommendation for achieving this balance in animal research.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →