← Latest papers
⚛️ high-energy experiments

Machine learning method for enforcing variable independence in background estimation with LHC data: ABCDisCoTEC

This paper introduces ABCDisCoTEC, an enhanced machine learning method that improves background estimation in LHC data by directly minimizing non-closure through a differentiable loss term and utilizing the modified differential method of multipliers to ensure robust variable independence, thereby overcoming the limitations of previous ABCDisCo approaches.

Original authors: CMS Collaboration

Published 2026-09-09
📖 6 min read🧠 Deep dive

Original authors: CMS Collaboration

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the high-stakes world of particle physics, scientists act as cosmic detectives, sifting through billions of tiny collisions to find a single, rare event that might reveal a new law of nature. These collisions happen inside the Large Hadron Collider, a massive ring buried beneath the Swiss-French border, where beams of protons smash together at nearly the speed of light. The resulting spray of particles is recorded by detectors like CMS, creating a mountain of data. The challenge is not just finding the new particle, known as the signal, but accurately knowing what the background noise looks like. If the background is estimated incorrectly, a scientist might mistake a common fluctuation for a discovery, or worse, miss a real discovery hidden in the noise. Traditionally, physicists have used a clever trick called the "ABCD method" to estimate this background. They divide their data into four boxes based on two different measurements. If the two measurements are completely unrelated to each other, the number of background events in the empty "signal box" can be predicted by multiplying the counts in the other three boxes. It is a reliable way to use real data to guess what the background looks like, but it has a fatal flaw: it only works if those two measurements are truly independent.

Finding two such independent measurements is incredibly difficult. In the complex chaos of a particle collision, almost everything is subtly connected. If the two measurements used to sort the data have even a tiny link between them, the prediction fails, and the background estimate becomes unreliable. For years, physicists have struggled to find the right pair of variables by hand, a process that is slow, difficult, and often unsuccessful when the signal is very similar to the background. A researcher from the CMS Collaboration at CERN has now developed a new solution that automates this difficult task using machine learning. They created a system that teaches a computer to invent its own two measurements, ensuring they are independent while still being excellent at spotting the signal. This new approach, which they call ABCDisCoTEC, allows them to estimate the background directly from the observed data with much higher precision, reducing the need to rely on imperfect computer simulations.

The core of this innovation lies in how the machine learning model is trained. In the past, researchers used a method called ABCDisCo, which taught a neural network to produce two outputs that were statistically independent. However, this method had a blind spot. It could ensure the two outputs were uncorrelated, but it could not guarantee that the background events would spread out smoothly across the data, which is a strict requirement for the ABCD method to work correctly. The new ABCDisCoTEC method fixes this by adding a specific instruction to the training process. The computer is now told to not only make the two outputs independent but also to minimize the "non-closure." In simple terms, non-closure is the error that occurs when the prediction from the three background boxes does not match the actual number of events found in the signal box. By directly teaching the network to reduce this error, the system learns to arrange the data in a way that makes the background estimation mathematically sound.

To test this idea, the researcher applied it to a search for a hypothetical particle called a "stealth supersymmetric top squark." This particle is a partner to the top quark, a heavy particle already known to exist, but it is predicted by theories that extend beyond our current understanding of physics. The problem with finding this specific particle is that it looks almost exactly like the background noise created by standard top quark pairs. In previous searches, the uncertainty in how well the background was modeled limited the ability to find a signal. The new method was put to the test using simulated data that mimics the conditions of the Large Hadron Collider. The results showed that the ABCDisCoTEC model successfully created two new variables that were independent of each other and greatly improved the stability and robustness of the background arrangement. This allowed the researcher to predict the background in the signal region with high accuracy, significantly improving the sensitivity of the search, with specific details and results from the search documented in a separate reference.

A major hurdle in using such complex machine learning models is tuning the many knobs and dials, known as hyperparameters, that control how the network learns. Usually, scientists have to guess these settings or run thousands of trials to find the right combination, a process that is time-consuming and prone to error. The paper also introduces a second innovation to solve this: a mathematical technique called the modified differential method of multipliers. Instead of guessing the right balance between the different goals of the training, this method treats the requirements as strict constraints. It forces the network to find a solution that meets specific targets for independence and background accuracy without needing to manually search through endless possibilities. This approach proved to be much faster and more stable than traditional methods, converging on the best solution in a fraction of the time.

The researcher then validated their findings by applying the method to real data collected from the Large Hadron Collider. They defined specific regions in the data, separate from the main search area, to check if the method worked correctly on actual collision events. The results showed that the behavior of the background in the real data matched the behavior seen in the simulations, with the non-closure error remaining small and consistent in the validation regions. This validation is crucial because it confirms that the method is robust enough to be used in real-world physics analyses, demonstrating that the machine learning model learned the underlying structure of the real data within these specific test regions.

The implications of this work extend far beyond this single search for a new particle. The ABCDisCoTEC method represents a significant step forward in automating the analysis of high-energy physics data. By allowing physicists to estimate backgrounds directly from observed data with high precision, it reduces the reliance on theoretical simulations, which can sometimes be inaccurate. This leads to more reliable measurements and a greater chance of spotting subtle signs of new physics. The combination of the new training method and the advanced optimization technique provides a powerful toolkit for future experiments. It demonstrates that machine learning can be used not just to classify data, but to fundamentally improve the statistical methods used to interpret the universe's most complex collisions. As the Large Hadron Collider continues to collect data, these tools will be essential in the ongoing quest to understand the fundamental building blocks of reality.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →