MadVfold: accelerating NLO event generation and reducing negative weights with SIMD vectorization and GPUs
This paper introduces MadVfold, a CUDACPP-based implementation of "vectorized folding" for MG5aMC that leverages SIMD and GPU acceleration to achieve 3x to 9x speedups in NLO event generation while significantly reducing negative weights.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the high-stakes world of particle physics, scientists at the Large Hadron Collider smash protons together to recreate the conditions of the early universe. To make sense of the debris, they rely on computer simulations that predict what should happen during these collisions. However, calculating these predictions with extreme precision is a monumental task. The most accurate methods, known as next-to-leading-order calculations, are incredibly slow and computationally expensive. A major hurdle in these simulations is the appearance of "negative weights," a mathematical quirk where some simulated events carry a negative value. To cancel out the noise these negative values create and achieve a clear picture, researchers must generate vastly larger batches of events than usual, consuming enormous amounts of computing power and time.
To solve this, physicists have developed a technique called "folding." Imagine trying to estimate the value of a complex integral by sampling a single point; it can be noisy and inaccurate. Folding is like calculating the same complex function for several different variations of a single event's kinematic variables and averaging the results. This approach provides a more stable estimate of the integral, which helps smooth out the statistical noise and reduces the need for negative weights. While effective, this method is itself very slow because it requires the computer to perform the same complex calculation many times for a single event. A new approach described by Andrea Valassi at CERN aims to speed up this process by using modern computer hardware to perform these repeated calculations simultaneously, rather than one by one.
Valassi's work introduces a method called "vectorized folding," which leverages the parallel processing power of modern graphics cards and advanced central processors. Instead of asking the computer to calculate the physics for one specific variation of an event, then the next, and so on, the new software groups dozens of these variations together. It then sends this entire batch to the computer's processor, which calculates all of them at the same time. This is a significant shift from the traditional way these simulations run, which processes events sequentially. By reorganizing the software to take advantage of this parallel capability, the researcher created a new tool called MADVFOLD, designed to work with the widely used Madgraph5_aMC@NLO simulation package.
The results of this reorganization are striking. In tests involving the collision of electrons and positrons to produce bottom quarks, the new method dramatically reduced the time required to generate simulated events. When using advanced processors found in modern data centers, the vectorized folding technique made the calculation roughly six to nine times faster than the standard method. When the calculations were offloaded to a graphics processing unit, the speedup was even more pronounced, reaching nearly ten times the speed of the original approach. These gains were achieved not by changing the underlying physics, but by changing how the computer executes the math, allowing it to handle the heavy lifting of folding much more efficiently.
The research also explored whether this parallel approach could be applied to generating events without the folding technique at all. Even in this "unfolded" scenario, where the goal is simply to speed up the standard simulation, the new software delivered a threefold increase in speed. This suggests that the architectural changes made to support folding have broader benefits, making the entire simulation pipeline more efficient. The work was developed using a novel process where the researcher relied heavily on large language models to assist in writing and testing code, a method that allowed the project to move from concept to working software in just around six weeks.
Despite the impressive speedups, the paper notes that folding remains a computationally expensive technique. Even with the new hardware acceleration, running a simulation with folding still takes significantly longer than running one without it. However, the trade-off is often necessary because folding reduces the number of negative-weight events, which in turn reduces the total number of events needed to reach a statistically significant result. By making the folding process itself much faster, the new method makes this necessary trade-off more manageable for future experiments, including the upcoming High-Luminosity LHC program. The study concludes that while the software is ready for further testing and refinement, it represents a significant step forward in making high-precision particle physics simulations more feasible for the massive datasets expected in the coming years.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.