Vision Transformers for End-to-End Quark-Gluon Jet Classification from Calorimeter Images
This paper presents a systematic evaluation demonstrating that Vision Transformer (ViT) and hybrid architectures outperform traditional Convolutional Neural Networks in classifying quark- and gluon-initiated jets from calorimeter images by effectively capturing long-range spatial correlations, thereby establishing new performance baselines for end-to-end deep learning in high-energy physics using public collider data.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Deep inside the massive ring of the Large Hadron Collider, particles smash together at nearly the speed of light, creating a chaotic spray of debris that physicists must sort through to understand the fundamental laws of the universe. When protons collide, they often produce two distinct types of particle sprays, known as jets: one born from a quark and the other from a gluon. While both look like messy cones of energy to a detector, they are fundamentally different. Quarks are the building blocks of protons and neutrons, while gluons are the glue that holds them together. Gluons carry a stronger "color charge," a property of the strong nuclear force, which causes them to break apart into a wider, more crowded spray of particles compared to the tighter, cleaner spray from a quark. Distinguishing between these two is crucial for scientists. If they cannot tell them apart, they might miss subtle signals of new, unknown physics hiding in the background noise, or they might mistake a common event for something extraordinary. For decades, physicists have relied on complex mathematical rules to separate these jets, but the patterns are so intricate and the data so vast that human-designed rules often fall short.
A team of researchers has now turned to a different kind of intelligence to solve this problem: artificial intelligence modeled after the human brain. Instead of trying to write equations that describe the difference between a quark and a gluon, they taught a computer to look at the raw images of the particle collisions and learn the difference on its own. They used data from the CMS experiment at CERN, which records the energy deposits left by particles as they pass through layers of detectors. These deposits are converted into digital pictures, where brighter spots represent more energy. The researchers tested a new type of artificial intelligence called a Vision Transformer. Unlike older computer vision tools that scan an image by looking at small, local patches one by one, a Vision Transformer looks at the entire image at once. It can instantly see how a bright spot in one corner of the image relates to a faint spot on the opposite side, capturing the global shape and structure of the jet in a single glance.
The study involved feeding the computer millions of these simulated jet images, half showing quark jets and half showing gluon jets. The images were constructed from three different types of detector data: the tracks left by charged particles, the energy measured in the electromagnetic calorimeter, and the energy measured in the hadronic calorimeter. By combining these into a single, multi-layered picture, the researchers gave the AI a complete view of the event. They then compared the performance of these new Vision Transformers against the traditional, long-standing tools used in the field, which rely on scanning small local areas. The results were clear and consistent. The Vision Transformers, especially when combined with other advanced architectures, proved significantly better at telling the two types of jets apart. They achieved higher accuracy and were better at correctly identifying the jets without making mistakes, a balance that is critical for scientific discovery.
The researchers found that the ability to see the whole picture at once gave the new models a distinct advantage. Quark and gluon jets differ in subtle, long-range patterns; the way energy spreads out across the entire detector is often more telling than the intensity of a single pixel. The older tools, which focus on local details, struggled to connect these distant dots. The Vision Transformers, however, thrived on this global context, successfully mapping the complex, sprawling structures of the gluon jets against the more compact quark jets. The study also explored different ways to mix these new models with older ones, creating hybrid systems that combined the best of both worlds. These hybrid models performed even better, suggesting that the future of particle physics analysis may lie in blending the local precision of traditional methods with the broad, contextual understanding of modern transformers.
While these findings are promising, the researchers are careful to note that their work was done using simulated data, a highly accurate computer model of how the detectors should behave. The next step will be to test these models on real data from the collider to see if they hold up under the messy, unpredictable conditions of a real experiment. Furthermore, these powerful models require significant computing power to run, which could be a hurdle for real-time analysis in the future. Despite these limitations, the study establishes a new benchmark for how machine learning can be applied to particle physics. It shows that by letting the computer see the entire event as a whole, rather than just its parts, scientists can unlock a deeper understanding of the subatomic world, potentially revealing new physics that has been hidden in plain sight all along.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.