← Latest papers
⚛️ high-energy experiments

Similarity Pairing with Energy Mover's Distance for Self-Supervised Pre-Training at the LHC

This paper introduces a data-driven, augmentation-free self-supervised pre-training method for the Large Hadron Collider that pairs distinct events based on their Energy Mover's Distance similarity to learn invariant representations, achieving downstream performance comparable to or better than traditional augmentation-based baselines while preserving event fidelity.

Original authors: Ho Fung Tsoi, Dylan Rankin

Published 2026-09-17
📖 4 min read🧠 Deep dive

Original authors: Ho Fung Tsoi, Dylan Rankin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the heart of Europe, the Large Hadron Collider smashes protons together at nearly the speed of light, creating a chaotic shower of new particles that physicists must sort through to understand the fundamental laws of nature. To make sense of this deluge of data, researchers rely on powerful computer models, often called foundation models, which are trained to recognize patterns and distinguish between different types of particle collisions. A major challenge in teaching these models is that they need to learn what stays the same when a particle collision is viewed from slightly different angles or with minor variations in how the detector records it. Traditionally, scientists have tried to teach this by artificially distorting the data, creating slightly altered copies of each event to show the model that the underlying physics remains unchanged despite the noise. However, in the high-stakes world of particle physics, these artificial distortions are risky; if a computer simulation changes a particle's path in a way that violates the laws of physics, the model learns the wrong lessons, becoming unreliable when faced with real data.

A team of researchers at the University of Pennsylvania has proposed a different way to teach these models, one that avoids artificial changes entirely. Instead of inventing fake variations of a particle collision, they simply look for two real, distinct collisions that happen to look very similar to one another. They use a mathematical tool called the energy mover's distance, which calculates exactly how much work it would take to rearrange the energy in one collision to match the layout of another. If the work required is small, the two events are considered close neighbors in the world of physics. By pairing these naturally similar events together, the researchers can train their model to recognize that these two different real-world occurrences share the same core identity, without ever having to break or bend the physical laws governing the data.

The researchers tested this idea using a massive dataset of simulated particle jets, which are sprays of particles created when quarks or gluons are knocked loose during a collision. They focused on two common types of jets, known as light quark jets and gluon jets, and used their new pairing method to train a computer model. Instead of taking a single jet and twisting it into a fake version, the system took two separate jets that were naturally close in their energy arrangement and asked the model to treat them as if they were the same thing. The model learned to ignore the tiny, random differences between these real events and focus on the deeper structure that defines them. When the researchers checked the results, the model had developed a clear mental map where different types of jets grouped together neatly, even for types of jets it had never seen during training.

To see if this new method actually worked for real-world problems, the team asked the model to act as a detective, looking for rare, unusual particles hidden among the common background noise. This is a critical task at the collider, where finding a new particle often means spotting a single odd event in a sea of millions of ordinary ones. The model trained with the natural pairing method proved to be just as good, and in some cases slightly better, at spotting these rare signals than models trained with the traditional method of artificial distortion. The results showed that by relying on the natural similarities found in real data rather than handcrafted tricks, the model learned a more robust and accurate understanding of particle physics. This approach suggests that the future of training these powerful tools may lie not in making up new data, but in finding the hidden connections that already exist within the vast amounts of real data we have collected.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →