← Latest papers
⚛️ phenomenology

Strong CWoLa: Binary Classification Without Background Simulation

This paper demonstrates that the Classification Without Labels (CWoLa) paradigm can eliminate the need for background simulation in high energy physics, enabling the training of supervised classifiers that achieve higher performance on real data by avoiding the domain shift caused by inaccuracies in low-level feature simulations.

Original authors: Samuel Klein, Matthew Leigh, Stephen Mulligan, Tobias Golling

Published 2026-08-27
📖 4 min read🧠 Deep dive

Original authors: Samuel Klein, Matthew Leigh, Stephen Mulligan, Tobias Golling

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the quest to understand the fundamental building blocks of the universe, physicists at the Large Hadron Collider smash particles together at nearly the speed of light. These collisions create a chaotic spray of new particles, some of which are fleeting and rare, while others are common and well-understood. The challenge for scientists is to find the needle in the haystack: identifying the rare, potentially new particles hidden within the overwhelming noise of ordinary debris. To do this, researchers rely on powerful computer programs called classifiers. These programs are trained to distinguish between the "signal" of a new discovery and the "background" of known physics. Traditionally, scientists have taught these programs using simulated data, where computers generate perfect examples of both the signal and the background. However, the real world is messy. The simulations, no matter how sophisticated, are never a perfect mirror of reality. As the data becomes more detailed and complex, the gap between the computer's simulation and the actual measurements grows wider, causing the trained programs to make mistakes when they face real data.

A team of researchers at the University of Geneva has developed a new way to train these classifiers that bypasses the need for simulated background data entirely. Their method, which they call strong CWOLA, allows a computer to learn how to spot a new particle by comparing a clean, simulated sample of that particle against a pile of real, unlabelled data from the collider. In this approach, the computer is not asked to tell the difference between two types of simulated events. Instead, it is asked to tell the difference between a dataset containing only the simulated signal and a dataset containing a mixture of real signal and real background. Because the real data contains the true, unfiltered behavior of the universe, the computer learns to recognize the signal based on how it actually appears in nature, rather than how it appears in a flawed model. This technique effectively removes the errors introduced by imperfect background simulations, allowing the classifier to perform better on real-world data than traditional methods.

The researchers tested this idea using data from a community challenge designed to mimic a search for a specific type of new physics: a heavy particle decaying into two jets of other particles. They created a scenario where the signal was a particle with a mass of 3.5 TeV, while the background consisted of ordinary particle collisions with a mass of 1.3 TeV. In their experiments, they trained classifiers using two different strategies. The first was the traditional approach, where the computer learned to distinguish the signal from a simulated background generated by a different physics program. The second was their new strong CWOLA method, where the computer learned to distinguish the simulated signal from a dataset of real collision data that contained a small, unknown amount of the signal mixed with the background. They tested this using both simple, high-level summaries of the data and complex, low-level details that describe the individual particles within the collision.

The results showed that the new method worked remarkably well. Classifiers trained with strong CWOLA performed just as well as, and in some cases better than, those trained on simulated background. This was true even when the real data used for training contained a significant amount of the signal itself, up to a level that would represent a major discovery in a real experiment. The researchers found that the method remained robust even when the signal was mixed into the background data at high rates, proving that the classifier could still learn the correct pattern without being confused by the contamination. When the team used the most detailed, low-level data available, the performance of all classifiers improved, but the strong CWOLA method maintained its advantage by avoiding the mismatches that occur when simulations fail to capture the full complexity of real particle interactions.

This work suggests that physicists can now train their most powerful tools without relying on the often-flawed simulations of background noise. By using the real data itself as a reference point, the computer learns a more accurate picture of what the signal looks like in the wild. The researchers demonstrated that this approach is not limited to simple cases but works effectively even when the data is complex and the signal is rare. While the study was conducted using simulated data to represent real-world conditions, the findings indicate that this technique could significantly improve the sensitivity of future searches for new physics. It offers a path forward where the limitations of computer modeling no longer hold back the ability to find the rarest and most elusive particles in the universe.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →