Functional anatomy of Pythia-Herwig differences with Kolmogorov-Arnold networks
This paper employs Kolmogorov-Arnold networks to perform a staged functional analysis of differences between Pythia and Herwig event generators, revealing how specific observable drivers like multiplicity and jet shape evolve and persist across shower, hadronization, and full-generator stages.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the high-energy world of particle physics, scientists do not see the fundamental collisions of nature directly. Instead, they rely on massive detectors that capture the debris of these events, recording signals that must be translated back into the language of theory. To make this translation, physicists use complex computer programs called event generators. These programs act as virtual laboratories, simulating how a collision between two protons evolves from a simple, high-speed impact into a shower of new particles that eventually form the jets and tracks seen in real detectors. Two of the most widely used programs are Pythia and Herwig. While both are built on the same fundamental laws of physics, they make different choices about how to model the messy, chaotic moments when particles interact and transform. Because these choices are not dictated by a single, perfect formula, the two programs often produce slightly different results. Understanding exactly where and why they differ is crucial; if the programs disagree, it is difficult to know whether a new discovery in the data is a sign of new physics or simply a quirk of the simulation.
A recent study by Arghya Chattopadhyay at the University of Puerto Rico at Mayagüez offers a new way to look at these disagreements. Rather than treating the difference between Pythia and Herwig as a single, blurry number or a global score, the researcher used a specialized type of artificial intelligence to dissect the problem into its smallest, most understandable parts. The goal was to trace the life of a single collision through the simulation, watching how the differences between the two programs change as the event evolves from a simple shower of particles to a fully formed jet of matter. By following the same starting point through three distinct stages of the simulation, the study reveals that the source of the disagreement is not fixed. Instead, the features that drive the programs apart shift dramatically depending on which stage of the process is being examined.
The researchers began by generating a set of identical high-energy collisions, creating a common starting point for both programs. They then let Pythia and Herwig evolve these collisions through three specific phases. In the first phase, the programs simulated only the initial shower of particles, stopping before any complex matter formation occurred. In the second phase, they added the process of hadronization, where the energetic particles clump together to form stable particles like protons and pions. In the final phase, they included the full complexity of the collision, accounting for the underlying spray of softer particles that accompanies the main event. Throughout this entire process, the team tracked a specific set of eight measurable properties of the resulting particle jets, such as how many pieces made up the jet, how heavy the jet was, and how spread out its energy was.
To understand the differences, the team employed a machine learning tool known as a Kolmogorov-Arnold network. Unlike standard neural networks that act as opaque black boxes, this specific type of network is designed to be transparent. It breaks down a complex prediction into a sum of simple, one-dimensional responses, one for each of the eight properties being measured. This allowed the researchers to isolate exactly which property was causing the programs to disagree at any given moment. They could then take the specific "answer" the network learned at one stage and test whether that same answer remained useful when applied to the later stages of the simulation.
The results showed a clear and surprising evolution in the nature of the disagreement. At the very first stage, when only the particle shower was active, the difference between Pythia and Herwig was driven almost entirely by the number of particles in the jet. The programs simply counted the pieces differently. However, once the process of hadronization was introduced, this multiplicity difference largely vanished. The programs began to agree on the number of particles, but they started to disagree on the mass and the shape of the jets. In the final, most complex stage, the disagreement became a mix of both shape and particle count, with the shape of the jet becoming the most significant factor.
Perhaps the most revealing finding was that information learned at the earliest stage did not always carry over to the later stages. When the researchers took the specific rule the network learned about particle counts at the shower stage and applied it to the later, hadronized events, it successfully reduced the disagreement between the programs. This suggested that the way the programs counted particles at the beginning had a lasting influence. However, when they tried to do the same thing with the shape of the jet, the rule learned at the shower stage failed to fix the disagreement later on. Even though the shape of the jet became very important in the final stages, the specific way the programs differed about shape at the beginning was not the same as how they differed at the end. The early shape difference simply did not persist.
The study also uncovered a limitation in how these simulations can be compared. For the mass of the jets, the network learned a very strong difference, but this difference was driven by only a tiny handful of rare events in the simulation. Because the data supporting this difference was so sparse, the researchers could not reliably use it to adjust the simulation. This highlighted a critical distinction: just because a computer model can identify a difference does not mean that difference is statistically robust enough to be used for correction.
This work provides a functional anatomy of the disagreement between two major tools of modern physics. It demonstrates that the gap between Pythia and Herwig is not a static error but a dynamic structure that reorganizes itself as the simulation progresses. The factors that matter most at the beginning of a collision are not necessarily the factors that matter most at the end. By using a transparent machine learning approach, the study allows physicists to see exactly which physical features are driving these differences and to understand how those features evolve. This level of detail is essential for ensuring that when scientists look for new physics in their data, they are not misled by the hidden assumptions of their own simulations.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.