Weighted Holm Procedures: Theory, Properties, and Recommendations
This paper systematically compares the weighted Holm procedure (WHP) and the weighted alternative Holm procedure (WAP), demonstrating through theoretical analysis, graphical representations, and simulations that WHP is uniformly more powerful, monotonic, and optimal for controlling the familywise error rate without violating its constraints, leading to a recommendation for its preferential use over WAP.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a judge presiding over a courtroom with 100 different cases (hypotheses) to decide on at once. You have a strict rule: you can only make one wrong decision (convicting an innocent person) in the entire batch, or the whole trial is considered a failure. This is what statisticians call controlling the "Familywise Error Rate" (FWER).
However, not all cases are created equal. Some are high-profile murder trials (very important), while others are minor parking tickets (less important). You want to be extra careful with the murder trials, but you also want to give the parking tickets a fair shot.
This paper is about two different ways to organize your courtroom to handle these cases fairly while keeping your "one mistake" rule intact. The authors, Beibei Li and Wenge Guo, compare two methods: WHP and WAP.
Here is the breakdown in simple terms:
The Two Methods: Sorting the Cases
Imagine you have a pile of evidence (p-values) for all 100 cases. You need to decide which ones to reject (convict) and which to keep.
1. The "Weighted" Method (WHP): The Priority Sort
- How it works: Before you even look at the evidence, you assign a "priority score" (a weight) to each case. A murder trial gets a high score; a parking ticket gets a low score.
- The Strategy: You mix the evidence with the priority score. You create a "weighted score" for every case. Then, you sort the cases from the lowest weighted score to the highest.
- The Logic: You tackle the most "important" cases first. If a murder trial has weak evidence, you might still reject it because its priority is so high. You are essentially saying, "Because this case matters so much, I will lower the bar slightly for it to get a closer look."
- The Result: This method is smarter and stronger. It finds more guilty parties (true discoveries) without making more mistakes.
2. The "Alternative" Method (WAP): The Raw Sort
- How it works: You ignore the priority scores when sorting. You look strictly at the raw evidence (the p-values) and sort the cases from the weakest evidence to the strongest, regardless of whether it's a murder trial or a parking ticket.
- The Strategy: You treat the sorting as "objective." You say, "The evidence speaks for itself."
- The Catch: Once you sort them, you apply the priority scores to decide how strict the bar is for each one.
- The Problem: This creates a weird disconnect. You might sort a low-priority parking ticket first because its evidence looks slightly better than a murder trial's, but then you apply a very strict rule to it because it's "low priority." It's like putting a VIP in the back of the line just because they arrived a few seconds later, then treating them like a VIP anyway. This inconsistency makes the method less powerful.
The Big Discovery: WHP Wins Every Time
The authors ran simulations (like running the courtroom 5,000 times with different scenarios) and found a clear winner:
- WHP (The Priority Sort) is uniformly more powerful. This means it catches more true effects (finds more guilty people) than WAP, without ever breaking the "one mistake" rule.
- WAP (The Raw Sort) is sometimes okay, but it often misses things that WHP would have caught. It's like using a net with bigger holes; you let some fish escape.
Why Does WHP Work Better? (The Analogy)
Think of the "weights" as magnifying glasses.
- WHP puts the magnifying glass on the most important cases before you look at them. It helps you see the details of the important cases immediately.
- WAP looks at everything with the naked eye first, decides what to look at, and then tries to use the magnifying glass. By the time you get to the important cases, you might have already missed the subtle clues because you were looking at the wrong things first.
The "Graphical" Picture
The paper also introduces a cool way to visualize this using a flowchart (a graph).
- Imagine a bucket of water (your total allowed error rate, ) being poured into different cups (the hypotheses).
- When a cup is "rejected" (a case is solved), the water from that cup flows into the remaining cups, making them easier to solve.
- WHP flows the water in a way that always helps the most important cups get filled up faster.
- WAP flows the water based on a different rule that sometimes leaves the important cups with less water than they could have had.
The Bottom Line for Real Life
If you are a researcher, a doctor, or a data analyst running a clinical trial (like testing a new drug):
- Use WHP (The Weighted Holm Procedure). It is the "Gold Standard." It is mathematically proven to be the best way to handle weighted hypotheses. It gives you the best chance of finding real effects while keeping your error rate safe.
- Avoid WAP unless you have a very specific, rare reason to do so. It is an older method that, while "objective" in how it sorts, ends up being less effective at finding the truth.
In short: Don't let the "objectivity" of raw numbers fool you. When you know some things are more important than others, let that importance guide your sorting process. WHP does exactly that, making it the superior choice for modern science.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.