A heterogeneous and vectorized sequence for the HL-LHC full tracking reconstruction of the CMS experiment
This paper outlines the new baseline strategy for CMS Phase-2 tracking at the HL-LHC, which combines GPU-optimized and CPU-vectorized algorithms to efficiently handle unprecedented computational challenges while reducing resource requirements and enhancing physics reach through displaced tracking and machine learning.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the heart of the European Alps, beneath a ring of superconducting magnets, the Large Hadron Collider smashes protons together at energies never before reached. The goal is to recreate the conditions of the universe just moments after its birth, searching for the fundamental rules that govern matter and energy. To do this, the CMS experiment acts as a massive, three-dimensional camera, capturing the debris of these collisions. However, the machine is about to undergo a transformation that will make it far more powerful, yet far more difficult to read. The High-Luminosity Large Hadron Collider will fire protons in bunches so dense that a single moment of collision will contain hundreds of overlapping events. Instead of seeing a clean picture of one interaction, the detector will be flooded with a chaotic storm of particles, creating a jumble of signals that must be untangled in real time.
The challenge is not just seeing the particles, but connecting the dots. As these particles fly through the detector, they leave behind tiny electrical signals, like footprints in the snow. The computer must link these footprints to reconstruct the path of each particle. In the past, this was a manageable task, but with the new machine, the number of footprints will increase by a factor of ten. If the software used to connect these dots remains the same, it will be overwhelmed, unable to keep up with the data stream. The result would be a loss of precious scientific information, as the computer simply cannot process the sheer volume of collisions.
To solve this, researchers at the CMS experiment have developed a new strategy for how the computer thinks about these tracks. They realized that the old way of working, which followed a slow, step-by-step process, was no longer viable. Instead, they built a system that works like a highly organized, parallel workforce, designed to run on the most modern computer chips available. This new approach combines three distinct methods, each optimized for a different type of processor, to sort through the chaos of the High-Luminosity collider. The result is a system that is not only faster but also significantly more efficient at filtering out incorrect paths while maintaining the same ability to find the rare, interesting particles hidden within the noise.
The new system begins by creating a list of potential starting points for particle paths. In the old method, the computer would generate a massive number of guesses, many of which were wrong, with a seeding fake rate as high as 80%. This was inefficient. The new strategy uses a specialized algorithm called Extended Patatrack to look at the innermost layers of the detector first. This tool is designed to run on graphics processing units, the same type of powerful chips found in video game consoles, which excel at doing many calculations at once. By requiring a minimum number of signals to form a starting point, this algorithm creates a very clean list of candidates, discarding the vast majority of false starts before they even begin. This reduces the workload for the next stage by a significant margin, cutting the number of false leads to less than 10% compared to the previous method.
However, some particles are tricky. Long-lived particles, which travel a short distance before decaying, might not leave enough signals in the inner layers to be caught by the first algorithm. To catch these, the team introduced a second tool called Line Segment Tracking. This method looks at the outer layers of the detector, where the geometry of the sensors allows it to build short segments of tracks in parallel. It then links these segments together, using geometric rules and machine learning to decide which ones fit together to form a complete path. This ensures that even the elusive particles that drift away from the center of the collision are not lost. When combined with the first tool, this two-pronged approach captures almost all the particles that need to be found, while keeping the number of mistakes extremely low.
Once the starting points are identified, the final step is to fit the full path of the particle through the detector. This is where the third tool, known as mkFit, comes in. This algorithm is designed to run efficiently on standard computer processors, using a technique called vectorization that allows it to process many tracks simultaneously. It takes the candidates from the previous steps and calculates their exact paths, determining their speed and direction with high precision. The team tested this entire new sequence in simulations that mimic the conditions of the future collider, where two hundred collisions happen at once. They found that the new system could reconstruct the tracks with equivalent efficiency to the old one, but with a much lower rate of errors.
The most striking result of this work is the speed. By moving the most demanding parts of the calculation to specialized hardware and optimizing the code for modern processors, the team reduced the time needed to process a single event. When running on standard computer processors alone, the new system is about nine percent faster than the old one. But when the heavy lifting is offloaded to graphics processors, the time required drops by thirty-three percent. This is a critical improvement, as the strict time limits of the online trigger system mean that every fraction of a second counts. The new system also improves the ability to find particles that have traveled a distance away from the collision point, a capability that is essential for studying certain types of rare physics.
This work represents a fundamental shift in how the experiment handles data. It moves away from a single, sequential process that struggles with complexity, toward a flexible, parallel approach that thrives on it. The algorithms developed for the online trigger are also being adapted for the offline analysis, where scientists study the data in greater detail after the fact. This means that the entire experiment will benefit from a unified, efficient way of seeing the universe. As the High-Luminosity Large Hadron Collider comes online, this new tracking sequence will ensure that the CMS experiment can continue to make precise measurements and discover new physics, even in the face of the most extreme conditions nature can provide.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.