← Latest papers
📊 statistics

The Elliptically Optimal Confidence Interval: A Bivariate Extension of Wilson's Score Method

This paper introduces the Elliptically Optimal (EO) confidence interval, a novel bivariate extension of Wilson's score method that derives explicit, non-degenerate bounds for the difference between two independent binomial proportions by projecting an elliptical joint region onto the estimand, thereby guaranteeing coverage without under-coverage while quantifying the trade-off of over-coverage in extreme cases.

Original authors: Nawaf Mohammed

Published 2026-09-11
📖 7 min read🧠 Deep dive

Original authors: Nawaf Mohammed

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of scientific measurement, researchers often need to compare two groups to see if they differ in a meaningful way. Imagine a doctor testing a new drug against an old one, or a biologist comparing the survival rates of two species of plants. The data usually comes in the form of simple counts: how many people got better, how many seeds sprouted. From these counts, scientists calculate a proportion, a percentage representing the success rate. But a single number is rarely enough to tell the whole story. Because the data comes from a sample rather than a census of the entire universe, there is always a margin of uncertainty. To communicate this uncertainty honestly, scientists construct a range of values, known as a confidence interval. This range is designed to capture the true difference between the two groups with a high degree of certainty, usually 95 percent. If the range is too narrow, it might miss the truth; if it is too wide, it becomes useless for decision-making. The challenge has long been finding the perfect balance: a range that is as tight as possible without ever falsely excluding the truth.

For decades, the standard tool for this job has been a method called the Wald interval. It is simple to calculate and appears in almost every introductory statistics textbook. However, this paper reveals that the Wald method has a hidden flaw. When researchers use it, especially with small sample sizes or when the success rates are very high or very low, the method tends to underestimate the uncertainty, producing intervals that are too narrow. While it often misses the true difference more often than the advertised 5 percent of the time, it does not do so uniformly; in specific cases, such as when sample sizes are small and success rates are very low, it can actually over-cover substantially, giving a false sense of security. To fix this, the author of this study, an independent researcher, has developed a new approach called the Elliptically Optimal confidence interval. This method does not try to guess the uncertainty; instead, it calculates the worst-case scenario for the uncertainty and builds the interval around that. The result is a range that is designed to never miss the truth under the normal approximation, though it sometimes becomes wider than necessary to ensure that safety.

The core of this new method lies in how it handles the unknown variables in the calculation. When comparing two groups, the uncertainty depends on the true success rates of both groups, which are unknown. The traditional approach is to plug in the observed rates from the sample to estimate this uncertainty. The new method takes a different path. It acknowledges that the true rates could be slightly different from the observed ones, and it asks: what is the largest possible uncertainty that could exist given the data we have? By maximizing the uncertainty rather than estimating it, the method creates a safety net. The author visualizes the possible combinations of the two groups' success rates as a shape on a graph. In the old methods, this shape is often treated as a simple point or a line. In this new approach, the shape is an ellipse, a stretched circle that represents all the plausible pairs of success rates that fit the data. The goal is to find the widest and narrowest possible difference between the two groups that still fits inside this elliptical shape.

Solving this problem required the author to map out the geometry of this ellipse and determine exactly where its edges lie. The ellipse is not a perfect circle; it is tilted and stretched, and it sits inside a square that represents the limits of probability, from zero to one hundred percent. The author discovered that the optimal interval is simply the range of differences covered by this ellipse. To make this practical, the author derived a set of rules that tell researchers exactly which formula to use based on the data they have. These rules cover six different scenarios, ensuring that the calculation is always correct, no matter how extreme the results are. Unlike the old methods, which can sometimes produce impossible results—such as a negative percentage or a range that extends beyond 100 percent—this new method always produces a valid, sensible answer. It never collapses into a single point, and it never leaves the realm of possibility.

The study also quantified exactly how much "extra" coverage this new method provides. Because it is designed to be safe, it is sometimes wider than strictly necessary. The author calculated that this extra width is not random; it follows a precise pattern. The interval is perfectly accurate when the two groups have certain specific relationships, but it becomes more conservative when both groups have very low or very high success rates. In those extreme cases, the interval widens significantly to ensure it does not miss the truth. This behavior is not a bug but a feature. The author showed that this conservatism is a fixed cost that does not disappear even if the sample size grows very large. This is a crucial distinction from other methods that might get better with more data but still fail in specific situations.

When the author compared this new method against the standard Wald interval and another popular alternative known as the Newcombe interval, the results were clear. The Wald interval consistently fell short of the 95 percent target in almost every situation tested, though it exhibited erratic behavior, occasionally over-covering in specific extreme cases. The Newcombe interval was the most accurate in the middle of the range, tracking the target very closely, but it did not offer a guaranteed safety net. The new Elliptically Optimal interval, while sometimes wider, provided a theoretical guarantee that the other methods could not under the normal approximation. However, the author noted that under the exact binomial distribution, the new method can still fall slightly below the nominal level in a small fraction of cases, though the shortfall is small. The author concludes that for situations where missing the truth is unacceptable—such as in regulatory approvals or critical medical decisions—this new method is the superior choice. It trades a bit of precision for a guarantee of safety. For situations where the goal is simply to get as close to the target as possible without a strict guarantee, the Newcombe method remains a strong contender. The paper does not claim to have solved the problem of statistics forever, but it has provided a rigorous, mathematically proven tool for those who need to be certain that their conclusions are not an illusion.

The study also highlights a fundamental truth about statistical estimation: there is no free lunch. You cannot have an interval that is both the shortest possible and guaranteed to never miss the truth. The old methods tried to be short and ended up missing the truth. The new method accepts that it must be longer to be safe. This trade-off is now visible and calculable. Researchers can look at their data and know exactly how much extra width they are paying for their safety. This transparency allows for better decision-making. If a researcher is in a situation where the success rates are expected to be extreme, they can anticipate that the interval will be wide and plan accordingly. If they are in a more moderate situation, the interval will be tighter. The method adapts to the data while maintaining its core promise.

In the end, this work is about restoring trust in the numbers. It shows that by thinking carefully about the geometry of uncertainty, rather than just plugging numbers into a formula, we can build better tools for science. The author has taken a problem that has been debated for decades and provided a solution that is both mathematically sound and practically useful. It is a reminder that in science, the most important thing is not just finding an answer, but knowing how sure you can be of that answer. The new interval provides that certainty, ensuring that when scientists say they are 95 percent confident, they truly mean it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →