← Latest papers
⚛️ high-energy experiments

Learning to bin: differentiable and Bayesian optimization for multi-dimensional discriminants in high-energy physics

This paper proposes a differentiable and Bayesian optimization framework for automatically determining optimal bin boundaries in multi-dimensional discriminants using Gaussian Mixture Models, demonstrating superior signal sensitivity compared to traditional hand-crafted or one-dimensional binning strategies in high-energy physics analyses.

Original authors: Johannes Erdmann, Nitish Kumar Kasaraguppe, Florian Mausolf

Published 2026-09-10
📖 4 min read🧠 Deep dive

Original authors: Johannes Erdmann, Nitish Kumar Kasaraguppe, Florian Mausolf

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the vast, high-speed collisions of modern particle physics, scientists are constantly hunting for the rare and the new. Imagine a camera taking billions of snapshots of subatomic particles smashing together, creating a chaotic spray of debris. To find a specific, fleeting particle hidden in this storm, researchers must sort the data into manageable groups. They use mathematical tools called discriminants to separate the interesting events, known as signals, from the overwhelming noise of ordinary background events. The standard way to do this sorting is to draw lines, or boundaries, on a graph to create boxes, or bins, where events are counted. If the boxes are too wide, the subtle signal gets lost in the noise; if they are too narrow or poorly placed, the statistical certainty of the measurement crumbles. For decades, physicists have drawn these boxes by hand, often using simple, evenly spaced lines or relying on the output of machine learning algorithms that force complex, multi-dimensional data into one-dimensional strips. This manual approach is practical, but it risks missing the most sensitive ways to organize the data.

A team of researchers at RWTH Aachen University in Germany has developed a new way to draw these boxes automatically, optimizing them to find the signal with the greatest possible clarity. Instead of guessing where the lines should go, they created a flexible system that learns the best shape for each category directly from the data. They treated the problem of sorting events as a puzzle where the goal is to maximize the chance of spotting a signal while keeping the background noise under control. To solve this, they built a model that can define bin boundaries in complex, multi-dimensional spaces, not just as simple straight lines but as curved, adaptive shapes that follow the natural structure of the data. They tested two different mathematical strategies to find these optimal shapes: one that uses a smooth, step-by-step calculation to adjust the boundaries, and another that uses a probabilistic method to explore different possibilities and learn from the results.

The researchers tested their methods using simulated data that mimics the kinds of signals and backgrounds found in real particle physics experiments. In these simulations, they created scenarios with one signal and five types of background noise, as well as scenarios with two different signals competing against the background. They compared their new, automated methods against the traditional way of sorting data, which involves taking a complex, multi-dimensional score and projecting it onto a single line to draw simple, evenly spaced boxes. In the simpler cases involving a single dimension, both of their new methods performed significantly better than the traditional, evenly spaced boxes, finding the signal with much higher sensitivity using the same number of categories. However, the real breakthrough appeared when they tackled the more difficult, multi-dimensional problems. Here, the traditional method of flattening the data onto a single line caused a loss of information, whereas their new approach could draw boundaries that curved through the multi-dimensional space, capturing the signal more effectively.

One of the key findings was that the method relying on smooth, step-by-step calculations, which the authors call a gradient-based approach, outperformed the probabilistic exploration method when the number of categories and the complexity of the data increased. This is because the probabilistic method struggles to find the best solution when there are too many variables to adjust at once. The researchers also showed that their system could be guided by practical rules. For instance, they could instruct the system to avoid creating categories that are too empty or have too much statistical uncertainty, ensuring that the resulting bins are robust and reliable for real-world analysis. In scenarios where the two signals were very similar and hard to tell apart, their multi-dimensional method provided a clear advantage, finding a level of sensitivity that the traditional, one-dimensional projections could not reach without using a vastly larger number of bins.

The team released their work as lightweight software tools that can be easily added to existing physics analysis workflows. These tools allow physicists to stop manually guessing where to draw their lines and instead let the data itself dictate the most effective way to categorize events. By moving away from rigid, hand-crafted boxes to flexible, learned boundaries, the researchers have provided a way to squeeze more information out of the same amount of data. This does not just improve the precision of current measurements but also offers a more powerful way to search for rare phenomena that might otherwise remain hidden in the noise. The work demonstrates that in the quest to understand the fundamental building blocks of the universe, the way we organize our observations is just as critical as the observations themselves.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →