Improved null proportion estimators for multiple discrete tests with plug-in FDR control
This paper introduces a general class of null proportion estimators and proposes new strategies that leverage null distribution information to reduce conservativeness and improve efficiency in multiple discrete hypothesis testing while maintaining valid plug-in FDR control.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a massive case involving hundreds of suspects (hypotheses). Your job is to figure out which ones are actually guilty (reject the null hypothesis) and which ones are innocent (keep the null hypothesis).
In the world of statistics, there's a famous rule called the Benjamini-Hochberg (BH) procedure. Think of this as a standard "guilt threshold." If a suspect's evidence (a p-value) is strong enough to pass this threshold, you declare them guilty. This rule is great because it controls the "False Discovery Rate" (FDR)—essentially, it promises that if you arrest 100 people, only a small, controlled percentage of them will actually be innocent.
However, there's a catch. The standard rule assumes that the "evidence" for innocent people looks like a perfectly smooth, continuous slide from 0 to 1 (like a ramp). But in the real world, especially when dealing with counts (like counting patients or genes), the evidence often looks like a staircase. You can only stand on specific steps, not anywhere in between.
The Problem: The "Staircase" Trap
When the evidence is a staircase (discrete data) instead of a smooth ramp (continuous data), the standard detective rule gets too scared. It becomes overly cautious.
- The Analogy: Imagine a security guard at a club who is supposed to let in 10% of the people who look suspicious. If the people are walking on a smooth ramp, the guard does a good job. But if they are walking on a staircase, the guard gets confused. Because the steps are "jumpy," the guard thinks, "This person is almost at the top step, but maybe they aren't quite there yet." To be safe, the guard lets in almost no one.
- The Result: The guard (the statistical test) becomes too conservative. They miss a lot of actual "guilty" suspects (low power) because they are terrified of accidentally letting in an innocent one.
The Solution: New Estimators
The authors of this paper, Iqraa Meah and Sebastian Döhler, say, "We can do better." They propose a new set of tools (estimators) to help the detective count how many suspects are actually innocent () more accurately, even when the evidence is on a staircase.
They introduce a "General Class" of tools that includes old, famous tools (like Storey's and Pounds-Cheng's) but fixes them for the staircase world. They offer three creative ways to fix the "over-cautious" problem:
1. The "Customized Ruler" (Rescaling)
- The Idea: The old tools used one giant ruler to measure everyone. But on a staircase, every step might be a different height.
- The Fix: Instead of using one ruler, they give each suspect their own custom ruler that fits their specific step size perfectly. This ensures the measurement is fair and doesn't overestimate how many people are innocent.
2. The "Mid-Step" Guess (Conditional Mean)
- The Idea: If you are standing on a step, the old tools assume you are at the very top edge of that step. But statistically, you might be somewhere in the middle.
- The Fix: They use a "mid-p-value" approach. Instead of guessing you are at the top of the step, they guess you are right in the middle of the step. This smooths out the staircase, making it look more like the smooth ramp the original rules were designed for. It's like filling in the gaps between the stairs with a ramp just for the calculation.
3. The "Magic Dice" (Expected Randomization)
- The Idea: Sometimes, the steps are so jagged that even the middle guess isn't perfect.
- The Fix: Imagine rolling a magic die for every suspect to see if they "slide" up or down slightly within their step. You do this thousands of times in a computer simulation and take the average result. This creates a "smooth" version of the jagged data.
- The Catch: While this is the most accurate mathematically, it requires a lot of computer power (like rolling dice a million times), so the authors suggest using it carefully.
What Did They Prove?
The paper claims three main things:
- Safety First: All these new methods still guarantee that the "False Discovery Rate" stays under control. The detective won't accidentally arrest too many innocent people.
- Better Efficiency: Because these methods stop being overly cautious, they catch more of the actual "guilty" suspects. They are more powerful.
- Fixing Old Tools: They showed that even the old Pounds-Cheng tool, which was previously unproven for this specific type of control, can be slightly tweaked to work perfectly with their new safety guarantees.
Real-World Testing
The authors tested these ideas in two ways:
- Simulations: They created fake data with "staircases" (discrete tests) and saw that their new methods found more "guilty" suspects without letting in more innocent ones compared to the old methods.
- Real Data: They applied this to real biological data (from the International Mouse Phenotyping Consortium), where they were looking at how gene changes affect mouse traits. The results were the same: the new methods worked better, finding more significant connections between genes and traits than the old, overly cautious methods.
Summary
In short, the paper says: "Stop treating jagged, step-like data as if it were smooth. By adjusting your measuring tools to fit the steps (using rescaling, mid-steps, or randomization), you can be less afraid of making mistakes, which allows you to find more true discoveries without breaking the rules of statistical safety."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.