← Latest papers
⚛️ high-energy experiments

Performance of heavy-flavour jet identification in the CMS high-level trigger in proton-proton collisions at s\sqrt{s} = 13.6 TeV

This paper presents the design, commissioning, and performance of new deep-learning-based heavy-flavour jet identification algorithms deployed in the CMS high-level trigger for 13.6 TeV proton-proton collisions, which significantly improved signal efficiency for key physics processes including Higgs boson production and decay.

Original authors: CMS Collaboration

Published 2026-08-20
📖 6 min read🧠 Deep dive

Original authors: CMS Collaboration

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the heart of the European Organization for Nuclear Research, known as CERN, a machine called the Large Hadron Collider smashes protons together at speeds approaching that of light. These collisions recreate the conditions that existed just fractions of a second after the universe began, producing a shower of new particles. Among the most important of these are heavy quarks, specifically the bottom and charm varieties. When these quarks are created, they quickly transform into heavier particles that travel a tiny distance before decaying. This slight delay leaves a distinct mark: a secondary point of origin that is slightly displaced from the main collision spot. Physicists call this a displaced vertex. Identifying jets—sprays of particles that originate from these heavy quarks—is essential for many experiments, particularly those searching for the Higgs boson, a particle that gives mass to others. The Higgs boson often decays into these heavy quarks, making the ability to spot them a critical skill for understanding the fundamental laws of nature. However, the universe also produces a vast number of jets from lighter particles, creating a noisy background that can easily hide the rare signals scientists are looking for.

The challenge for the researchers at CERN is not just finding these heavy quarks, but finding them fast enough to keep up with the machine. The Large Hadron Collider produces collisions at a rate of 40 million times per second. It is impossible to save data from every single event; the storage systems simply cannot handle that volume. Instead, a sophisticated filtering system, called a trigger, must decide in a fraction of a microsecond which collisions are interesting enough to keep and which should be discarded. This system acts as a gatekeeper, reducing the massive flood of data down to a manageable few thousand events per second. For years, this gatekeeper relied on established methods to distinguish between jets from heavy quarks and those from lighter particles. But as the collider became more powerful, producing more collisions and more background noise, the old methods began to struggle. They were not sensitive enough to catch the rare, interesting events without letting too much noise through, or they had to be set so strictly that they missed valuable data.

To solve this, the CMS collaboration, one of the two main experiments at the collider, developed a new generation of tools for their high-level trigger system. This system runs on a farm of standard computer processors and is responsible for the final, most detailed selection of events. The team introduced a new type of algorithm based on deep learning, a form of artificial intelligence that mimics the way the human brain learns to recognize patterns. Specifically, they utilized a technique called a dynamic graph convolutional neural network. In simple terms, this algorithm treats the particles inside a jet not as a simple list, but as a connected web or graph. It looks at how every particle relates to its neighbors, analyzing the complex structure of the spray to determine its origin. This approach allowed the system to see subtle differences between a jet made of heavy quarks and one made of lighter particles that previous methods could not detect.

The researchers tested these new algorithms using data collected from proton-proton collisions at an energy level of 13.6 tera-electronvolts, a record high for the machine. They focused on two main tasks: identifying jets from bottom quarks and jets from charm quarks. The results showed a significant improvement. The new system could identify bottom quark jets with much higher accuracy while simultaneously rejecting the background of light jets far more effectively than the previous generation of tools. This meant that the trigger could be set to be more inclusive, keeping more of the rare events that scientists want to study without being overwhelmed by noise. Perhaps even more importantly, the new system made it possible to identify charm quark jets with high efficiency for the first time in this online filtering system. This opened the door to studying processes involving charm quarks that were previously too difficult to isolate in real-time.

The impact of these improvements was immediate and measurable across several key areas of physics. One major success was in the search for pairs of Higgs bosons decaying into four bottom quarks. This is a rare process that helps physicists understand how the Higgs boson interacts with itself. The new triggers allowed the experiment to collect data on these events with much greater efficiency, effectively doubling the number of useful events captured in certain energy ranges compared to the previous setup. Similarly, the system improved the ability to study the Higgs boson when it is produced alongside a pair of top quarks, another crucial channel for understanding the particle's properties. The new tools also enabled a dedicated search for the Higgs boson decaying into charm quarks, a process that had been nearly impossible to trigger on in the past due to the overwhelming background. By lowering the energy thresholds required to save an event, the new system allowed physicists to explore lower energy regions that were previously inaccessible.

The team also developed a specialized version of this algorithm to handle high-energy particles that decay so quickly their products merge into a single, large spray. This is common when very heavy particles are produced and travel at high speeds. The new system could identify these merged jets with high precision, further expanding the reach of the experiment. Throughout the data-taking period, the researchers continuously monitored the performance of these algorithms, comparing their predictions against the actual data collected. They found that the computer simulations used to design the system matched the real-world data very closely, confirming that the new tools were reliable and stable over time. The system maintained its high performance even as the number of simultaneous collisions in the machine increased, a condition known as pileup, which often confuses simpler detectors.

By the end of the data collection period, the new triggers had become the standard for the experiment. They were adopted for the analysis of hundreds of terabytes of data, contributing directly to improved sensitivity in searches for new physics. The work demonstrated that advanced machine learning techniques could be successfully deployed in the split-second environment of a particle physics trigger, a feat that was once thought too computationally demanding. The success of this project means that the CMS experiment can now see deeper into the data, finding signals that were previously hidden in the noise. This capability is essential for the future of the field, as the collider continues to operate at higher intensities, pushing the boundaries of what is known about the fundamental building blocks of the universe. The ability to distinguish the heavy from the light, the rare from the common, in real-time, has become a cornerstone of modern particle physics, ensuring that no potential discovery is lost to the limitations of the filter.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →