The Search Budget of the BSM Resonance Program
This paper quantifies the statistical trials factor of the ATLAS BSM resonance program by calculating that its current published record requires a local significance of 6.55σ for a 5σ global discovery, while a fully combinatorial scan would only marginally increase this requirement to 7.11σ, and further argues that two-stage unblinding is a superior safeguard against spurious signals in machine-learning searches where trials factors are difficult to enumerate.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the vast, high-energy collisions of the Large Hadron Collider, physicists are hunting for new particles that could rewrite the laws of nature. These searches often involve looking for a "bump"—a small, unexpected rise in the number of events at a specific mass—hidden within a smooth, predictable background of ordinary particle interactions. The challenge is that nature produces a staggering variety of these bumps, and researchers must scan a massive landscape of possible masses and particle combinations to find them. The more places they look, the higher the statistical bar they must clear to claim a discovery, because random fluctuations can mimic a signal if one looks hard enough in enough different spots. This is known as the "look-elsewhere effect," and it acts as a natural filter, ensuring that only the most robust signals are accepted as real discoveries.
A team of researchers at the University of Geneva has now mapped out exactly how much ground the ATLAS experiment has covered in its search for these new resonances, and what it would cost to scan every possible corner of the data. By counting the number of independent places the experiment has looked, they calculated the precise statistical threshold required to turn a local hint into a global discovery. Their analysis reveals that the current published searches, which cover a wide range of theoretical models, have already looked in roughly 7,900 distinct places. To claim a discovery today based on this existing record, a signal would need to be strong enough to correspond to a local significance of 6.55 standard deviations, rather than the traditional 5.
The study then asked a bolder question: what if the experiment scanned every single mass that could be built from up to four reconstructed particles, without being guided by specific theories? This hypothetical "combinatorial" scan would cover a much wider territory, amounting to 360,000 distinct looks. While this is a massive increase in the number of places to check, the statistical penalty is surprisingly small. The extra effort required to clear the bar for a discovery in this fully open scan would only raise the local significance requirement from 6.55 to 7.11. In other words, scanning 46 times more territory costs less than half a standard deviation in sensitivity. This finding suggests that the cost of being completely open-minded is far lower than many physicists feared, and that the program is not yet limited by the sheer number of places it has looked, but rather by the number of events available to fill the histograms.
However, the paper also identifies a new kind of risk that arises when using advanced machine-learning tools to find these bumps. Unlike traditional scans where the number of places looked is known and countable, some modern algorithms act as imperfect estimators that can occasionally produce false alarms. The authors found that if a machine-learning network flags a signal, the rate of these false alarms can be much higher than standard statistics would predict, effectively inflating the number of "looks" in a way that cannot be counted. To solve this, they propose a two-stage safety procedure. In the first stage, the algorithm scans a small fraction of the data to identify potential candidates. Only those specific candidates are then "unblinded" and checked against the remaining, unseen data. This method acts as a filter, ensuring that any signal that does not repeat in the second stage is discarded as a fluke. Their simulations show that this two-step approach is actually more sensitive than trying to correct for the errors in a single, massive scan, provided the machine learning tool has a known rate of mistakes.
The researchers also examined the specific landscape of known models to see how much of the territory remains unexplored. They found that while the published ATLAS program has scanned the vast majority of mass spectra motivated by current theories, there are still 51 specific combinations of particles that have been predicted by models but never searched for. These unscanned areas represent a small but distinct gap in the current coverage. The study concludes that the "search budget" for the resonance program is well-defined and manageable. The current published record is consistent with the number of random fluctuations expected, and the path forward involves either filling in the few remaining theoretical gaps or embracing the fully combinatorial scan, which remains statistically feasible and offers a powerful way to find the unexpected without needing to guess where to look.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.