← Latest papers
⚛️ phenomenology

Know What You Don't Flow

This paper demonstrates how heteroscedastic and Bayesian normalizing flows can learn and propagate calibrated systematic and statistical uncertainties for generative neural networks in LHC physics, using both toy models with explicit likelihoods and classifier-reweighted approaches for top pair events.

Original authors: Anja Butter, Sascha Diefenbacher, Tilman Plehn, Lorenz Vogel

Published 2026-08-25
📖 5 min read🧠 Deep dive

Original authors: Anja Butter, Sascha Diefenbacher, Tilman Plehn, Lorenz Vogel

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the heart of Europe, the Large Hadron Collider smashes protons together at speeds close to light, creating a shower of new particles that fly out in every direction. Physicists have spent decades mapping these collisions, turning the chaotic spray of debris into precise measurements of how the universe works. To make sense of this data, they rely on two things: the actual measurements from the detectors and the theoretical predictions from complex computer simulations. For years, these two have matched up beautifully, but as the collider becomes more powerful and collects more data, the margin for error shrinks. The measurements are becoming so precise that the theoretical predictions must be equally sharp. If the computer models are slightly off, or if the scientists cannot tell exactly how much they might be off, the entire search for new physics could be misled. The challenge is no longer just about getting the right answer, but about knowing exactly how confident we can be in that answer.

This is where a new approach from a team at the University of Heidelberg comes in. They are working on a way to teach artificial intelligence to not only predict the outcome of these particle collisions but also to admit when it is unsure. In the world of machine learning, a common tool called a generative network acts like a master forger. It learns the patterns of real particle collisions and then creates new, fake events that look indistinguishable from the real thing. This is incredibly useful because simulating real collisions is slow and expensive, while the AI can generate millions of fake events in seconds. However, until now, these forgers have been silent about their own mistakes. They would produce a result without telling the physicist whether the answer was rock-solid or a wild guess. The Heidelberg team has developed a method to fix this, giving these networks a voice to express their uncertainty.

The researchers focused on two distinct types of uncertainty that can plague these simulations. The first is statistical uncertainty, which happens simply because the training data is limited. If a network has only seen a few examples of a rare event, it should be less confident when predicting that event again. The second type is systematic uncertainty, which is more subtle. This arises from flaws in the network itself, the way it was trained, or the noise in the data it learned from. Even with infinite data, a network might still have a built-in bias or a blind spot. The team wanted to know if they could teach a single AI to track both of these sources of error simultaneously and report them accurately.

To test their idea, they started with a simple, made-up model where they knew the exact truth. They injected artificial noise into the data to simulate real-world imperfections and then watched how the AI reacted. They found that they could train the network to learn the size of the noise it was seeing. When they added more noise, the network's reported uncertainty grew larger, matching the reality of the situation. They also tested how the network handled the speed at which it learned. If the learning process was too fast and chaotic, the network developed a specific kind of error, and the team's method successfully identified this as a systematic flaw. Crucially, they showed that the network could distinguish between errors caused by a lack of data and errors caused by the training process itself, assigning the correct type of uncertainty to each.

The real test came when they applied this method to a complex, realistic scenario: the production of top quark pairs, one of the heaviest and most studied particles in the collider. In this case, they did not know the exact mathematical truth beforehand, which is the usual situation for physicists. To solve this, they used a clever trick involving a second AI, a classifier, to act as a guide. This classifier compared the rough, initial guesses of the generative network against the real data and helped correct the network's path. By using this corrected path as a target, the generative network could learn to produce accurate events and, at the same time, learn to quantify its own systematic errors. The results were striking. The network successfully mapped out where it was confident and where it was not, propagating these uncertainties through every direction of the particle collision.

The team also explored how these networks handle "nuisance parameters," which are variables in the simulation that are not perfectly known, such as the exact mass of the top quark. They trained the network to vary this mass and observed how the uncertainty in the final results changed. They discovered a clear, predictable relationship: as the uncertainty in the input mass grew, the uncertainty in the output events grew in a specific, measurable way. This means that in the future, physicists could use these networks to instantly see how a small change in a fundamental constant would ripple through their entire analysis.

The work demonstrates that it is possible to build generative networks that are not just fast forgers, but honest partners in scientific discovery. By teaching these systems to calibrate their own confidence, the researchers have provided a tool that can handle the extreme precision required by the next generation of particle physics experiments. The networks now do more than just generate data; they tell the scientists exactly how much they can trust that data, turning a black box of artificial intelligence into a transparent instrument for exploring the fundamental laws of nature.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →