The Benjamini--Hochberg Procedure Can Fail to Control the FDR for Correlated Two-Sided Gaussian Tests
This paper disproves a long-standing conjecture by demonstrating that the Benjamini–Hochberg procedure fails to control the false discovery rate at its nominal level for correlated two-sided Gaussian tests, a result rigorously proven via interval arithmetic and verified by the author.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to catch a group of liars in a room full of honest people. You have a special rulebook (the Benjamini–Hochberg procedure) that tells you how to shout "Gotcha!" without accidentally accusing too many innocent bystanders. For decades, statisticians believed this rulebook was unbreakable. They thought that even if the people in the room were whispering secrets to each other (correlated data), the rulebook would still keep the number of false accusations below a strict limit, say 1% (0.01).
But a new paper, written with the help of a super-smart AI named GPT-5.6 Pro, has just pulled the rug out from under that belief.
The Big Discovery: The Rulebook Has a Loophole
The authors show that for a specific type of tricky situation—where the "tests" are like two-sided Gaussian numbers (think of them as measurements that can wiggle up or down from zero)—the rulebook can actually fail. It turns out that when these measurements are correlated in a very specific, sneaky way, the rulebook lets the false accusation rate creep up just a tiny bit above the limit.
Instead of staying safely at 1%, the rate of false alarms can jump to 0.0104. That's a small number, but in the world of strict math, it's a giant crack in the dam. The paper doesn't just guess this; it builds a rigorous "certificate" (like a mathematical safety inspector's report) that proves, with absolute certainty, that the rate must be higher than 0.01 for large groups of tests.
The "Sneaky Factor" Model
How did they find this loophole? They built a fictional world, a "factor model," to test the rulebook. Imagine a giant room with three groups of people:
- The Honest Crowd (96% of the room): These are the true nulls. They are supposed to be innocent.
- The Signal Group A (1% of the room): These are the "real" signals, moving in one direction.
- The Signal Group B (3% of the room): These are also real signals, but moving in a different direction.
Here's the twist: Everyone in the room is listening to a single, invisible conductor (a "latent factor" called Z).
- The Honest Crowd moves with the conductor.
- The Signal Groups move against the conductor, but at different speeds.
When the conductor waves their baton in a certain way, the signals get louder, making the rulebook shout "Gotcha!" more often. But at the exact same time, the "Honest Crowd" gets a bit jittery, making their tails (the edges of their behavior) look suspiciously like the signals. This creates a perfect storm where the rulebook gets confused, rejects too many innocent people, and the false discovery rate spikes just enough to break the 1% rule.
How Sure Are They?
This isn't a "maybe" or a "we think." The authors are incredibly sure.
- The Proof: The core argument was generated by an AI and then carefully checked by the human author. They used a method called "interval arithmetic," which is like doing math with a safety margin that guarantees the answer is never wrong, even with rounding errors. They proved that for a large enough number of tests, the rate is strictly greater than 0.0104.
- The Simulation: To back this up, they ran computer experiments (Monte Carlo simulations). When they tested with 200 groups of data (totaling 20,000 tests), the computer measured the false discovery rate at 0.010359. This is statistically higher than the 1% limit, confirming the theory in a real-world simulation.
What This Means (and What It Doesn't)
This paper explicitly rules out the idea that the Benjamini–Hochberg procedure is safe for all correlated two-sided Gaussian tests. For twenty years, experts believed it was safe; this paper says, "Actually, it's not."
However, the paper is careful not to say the rulebook is useless. The violation is tiny (0.0104 vs 0.01), and it happens in a very specific, constructed scenario with a large number of tests. The authors admit they don't know if this happens with smaller numbers of tests or if there's a universal cap on how bad it can get. They also didn't test this on real-world data like genes or stocks yet; they built a mathematical model to prove the point.
So, the next time you hear that a statistical method is "proven" to work under any condition, remember this paper: sometimes, the most trusted tools have a tiny, hidden crack that only a super-smart detective (and a bit of AI help) can find.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.