← Latest papers
📊 statistics

Everywhere Valid Bounds on False Discovery Proportions in Conformal Inference

This paper introduces a finite-sample, distribution-free framework that establishes simultaneous, high-probability upper bounds on the false discovery proportion for all possible rejection thresholds in conformal inference, enabling flexible post hoc selection while providing tighter guarantees than existing methods.

Original authors: Ziang Song, Ying Jin, Emmanuel J. Candès

Published 2026-05-21
📖 5 min read🧠 Deep dive

Original authors: Ziang Song, Ying Jin, Emmanuel J. Candès

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to find a few rare, special items hidden in a massive warehouse. You have a "detector" (a machine learning model) that gives every item a "suspicion score." The lower the score, the more likely the item is a "fake" or an "outlier" (like a counterfeit bill or a defective part).

To catch the fakes, you set a threshold. If an item's score is below this line, you flag it for inspection.

  • The Problem: If you set the line too low, you miss the fakes. If you set it too high, you waste time inspecting real items that aren't fakes.
  • The Old Way: Traditionally, statisticians would say, "On average, if you run this test a thousand times, you'll only make a mistake 5% of the time."
    • The Flaw: This is like saying, "On average, your car won't break down." But if you are driving right now, that average doesn't help you. You need to know: "Is this specific car safe right now?" Also, if you look at your results and decide to change the threshold because you didn't find enough fakes, the old math breaks down completely.

This paper introduces a new, super-reliable safety net that works everywhere and all at once.

The Core Idea: The "All-Seeing" Safety Net

The authors created a method that draws a "ceiling" over your results. Imagine you are walking through a foggy forest (your data). You want to know how many dangerous animals (false alarms) are in the area.

Instead of guessing the average number of animals, the paper's method builds a transparent, high-probability fence that sits above the actual number of dangerous animals for every possible path you could take through the forest.

  • Simultaneous Validity: The fence doesn't just work for one specific path (one specific threshold). It works for every path you could possibly choose, even if you change your mind and pick a new path after looking at the forest.
  • The Guarantee: The paper proves that with 90% (or 95%) confidence, the actual number of mistakes you make will always be below this fence.

How They Built the Fence (The "Magic Trick")

The secret sauce is how they calculate this fence.

  1. The "Null" Universe: They realized that if the items you are flagging are actually "normal" (not fakes), their suspicion scores follow a very specific, predictable pattern, like a deck of cards shuffled perfectly.
  2. The Simulation: Instead of trying to guess the pattern using complex math formulas (which often result in a very loose, conservative fence), they simply simulated millions of "what-if" scenarios. They shuffled the "normal" cards over and over to see how high the number of false alarms could possibly jump.
  3. The Shape-Shifting Fence: Old methods used a straight, flat ceiling (a linear line). But the paper noticed that the "noise" in the data isn't flat; it's wiggly. Near the edges (very low or very high thresholds), the noise behaves differently.
    • Their new fence is flexible. It pinches in where the data is stable and puffs out where the data is wild. This makes the fence much tighter and more useful than the old, bulky straight lines.

Real-World Examples from the Paper

The authors tested this on two specific scenarios:

1. The Outlier Detective (Fraud Detection)

  • Scenario: You have a pile of credit card transactions. Most are normal; a few are fraud. You want to flag the fraud.
  • The Win: The new method tells you, "No matter what threshold you pick to flag fraud, we guarantee that at least 90% of the time, the percentage of innocent people you wrongly accused will be below this line." This is crucial because accusing an innocent person is expensive and stressful.

2. The Drug Hunter (Drug Discovery)

  • Scenario: A pharmaceutical company has a library of thousands of chemical compounds. They want to find the few that might cure a disease. They use a computer model to rank them.
  • The Win: The researchers want to pick the "top" candidates to test in a lab (which is expensive). The new method allows them to say, "If we pick the top 100 drugs based on this score, we are 90% sure that at least X% of them are actually promising."
  • The "Post-Hoc" Magic: In the past, if a researcher looked at the data, saw they only found 2 drugs, and decided to "relax the rules" to find 10, the old math would lie to them. This new method says, "It doesn't matter if you change the rules after looking at the data; the safety fence still holds."

Why This Matters

  • No More "On Average": It gives you a guarantee for the specific dataset you are holding in your hands, not just a theoretical average.
  • Freedom to Explore: Scientists can look at their data, adjust their criteria, and explore different thresholds without fear of breaking the statistical rules.
  • Tighter Bounds: Because the fence adapts to the shape of the data, it isn't as overly cautious (conservative) as previous methods, allowing researchers to find more true discoveries without increasing the risk of mistakes.

In short, the paper provides a universal, flexible, and mathematically proven safety net that ensures you don't accidentally flag too many innocent items, no matter how you choose to look at your data.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →