Mirror and knockoff+ thresholds under dependence
This paper demonstrates that mirror and knockoff+ thresholds can fail to control the false discovery rate under dependence, even when p-values are uniform and satisfy positive regression dependence, because their validity relies on the specific joint behavior of null signs rather than just marginal symmetry or standard distributional assumptions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a massive case with hundreds of suspects. You have a special test that gives each suspect a score: a high positive score means they are likely guilty, while a high negative score means they are likely innocent. The problem is that your test isn't perfect; sometimes innocent people get high scores just by bad luck. To be fair, you need to figure out how many of your "guilty" verdicts are actually mistakes. This is the world of multiple testing, a branch of statistics used everywhere from medical trials to genetics. The goal is to control the False Discovery Rate (FDR), which is simply the percentage of your "guilty" list that is actually made up of innocent people.
For decades, statisticians have used a clever trick to estimate these mistakes. They assume that if a suspect is truly innocent, their score is just as likely to be positive as it is to be negative, like a fair coin flip. So, if you see 100 people with high positive scores, you look at how many innocent people have high negative scores. If you see 10 innocent people with high negative scores, you guess that about 10 of your 100 "guilty" people are actually innocent. This "mirror" idea works beautifully when the suspects are independent—when one person's luck doesn't affect another's. But what happens if the suspects are connected? What if they all share a common source of noise, like a group of friends who all get the same bad weather report? This is the question that drives the research in this paper.
The paper, titled "Mirror and knockoff+ thresholds under dependence," asks a simple but startling question: What happens if we use this mirror trick on a group of suspects who are secretly influencing each other? The author, Xianyang Zhang, discovers that the method can completely fall apart, leading to a disaster where the number of innocent people wrongly accused is far higher than anyone expected.
The paper proves that the "mirror" method relies on a very specific, strong condition: that once you know how "loud" a suspect's score is, their sign (positive or negative) should be a fresh, independent coin flip. The paper shows that if this condition is missing—even if the scores look perfectly normal and symmetric on their own—the method fails.
Here are the three ways the paper shows this failure happens, using vivid scenarios:
1. The "Secret Mood" Trap (Non-Gaussian Example)
Imagine a group of eleven suspects. Unbeknownst to you, there is a "secret mood" variable that flips a coin. If the coin is heads, the secret mood makes it slightly more likely for all eleven suspects to get positive scores. If it's tails, they are more likely to get negative scores. Individually, every suspect still looks perfectly fair (their scores are uniformly distributed). However, because they are all reacting to the same secret mood, they tend to move together.
The paper calculates that if you use the mirror method on this group, the error rate jumps from the target of 10% to 17.4%. In a slightly different setup within the same family of examples, the error rate can get as high as 50%. This means half of your "guilty" list could be innocent, even though every single suspect looked fair on their own.
2. The "Common Wind" Problem (Fixed Gaussian Correlation)
Now, imagine the suspects are all standing in a field. There is a gentle, constant wind blowing from the east (a common factor). Even if the wind is very weak, it pushes everyone's score slightly to the right (positive).
The paper proves that if you have a huge number of suspects (hundreds or thousands), this tiny, shared wind causes a massive problem. As the group gets bigger, the mirror method starts to see way more positive scores than negative ones, not because the suspects are guilty, but because the wind pushed them all together. The paper shows that for any fixed positive correlation, no matter how small, the error rate eventually climbs to 50% as the number of hypotheses grows. This is a direct warning: even a tiny bit of shared noise can break the mirror method in large groups.
3. The "Different Scales" Nightmare (Worst-Case Scenario)
Finally, the paper looks at a scenario where the suspects have different "sensitivities." Some are very sensitive to the wind (large scale), while others are less sensitive (small scale). The author constructs a mathematical model where the wind blows in a way that tricks the mirror method perfectly.
In this worst-case scenario, the paper proves that the error rate can be made arbitrarily close to 100%. This means you could end up accusing a hundred innocent people and finding zero controls to stop you, simply because the scales of their scores were different.
What This Means
The paper is careful to say that this does not mean the famous "Knockoff" method (a more advanced version of the mirror trick) is broken. The advanced method works because it builds special "knockoff" controls that are designed to swap places with the suspects, preserving the independence of the coin flips. The problem arises when people try to use the simple mirror formula on data that doesn't have this special structure.
The main takeaway is that symmetry isn't enough. Just because every individual suspect looks fair doesn't mean the group acts fairly together. If the suspects are linked by a common factor, the mirror method can be fooled into thinking there are fewer mistakes than there really are. The paper uses exact mathematical proofs and computer simulations to show that these failures aren't just theoretical; they can happen with moderate group sizes (like 11 to 1,000 people) and are visible in real-world data.
In short, the paper warns us: Don't trust the mirror if the suspects are holding hands. If they share a common source of influence, the simple trick of counting negative scores to estimate positive mistakes will fail, potentially leading to a flood of false accusations.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.