← Latest papers
⚛️ high-energy experiments

Self-Supervised Learning for Robust Resonance Mass Regression in Cascade Decays

This paper proposes a self-supervised learning framework using VICReg and transformer encoders to pre-train models on corrupted data, demonstrating that this approach yields more robust and accurate heavy resonance mass reconstruction in complex cascade decays compared to traditional supervised training.

Original authors: Ho Fung Tsoi, Alex Yang, Luis Felipe Gutierrez Zagazeta, Shion Chen, Dylan Rankin

Published 2026-09-17
📖 4 min read🧠 Deep dive

Original authors: Ho Fung Tsoi, Alex Yang, Luis Felipe Gutierrez Zagazeta, Shion Chen, Dylan Rankin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the vast, high-speed laboratories where physicists smash particles together to understand the universe, the ultimate goal is often to find something that shouldn't exist. These experiments, conducted at facilities like the Large Hadron Collider, are essentially massive, ultra-sensitive cameras designed to capture the fleeting debris of collisions. When two particles collide at nearly the speed of light, they can briefly form a heavy, unstable particle that instantly shatters into a spray of smaller, lighter fragments. By measuring the energy and direction of these fragments, scientists can work backward to calculate the mass of the original, invisible particle. This calculation is the key to spotting new physics, but it is a fragile process. The detectors that record these collisions are not perfect; they sometimes miss pieces of the puzzle, miscalculate the energy of a fragment, or confuse one type of particle for another. These small errors, known as systematic uncertainties, can blur the final picture, turning a sharp, distinct signal of new physics into a vague, unrecognizable smear.

A team of researchers at the University of Pennsylvania and Kyoto University has developed a new way to sharpen this picture, even when the data is messy. They tackled a specific challenge: how to accurately measure the mass of a heavy, hypothetical particle that decays into a complex cloud of eleven different pieces, some of which escape detection entirely. In the past, scientists relied on computer models trained with labeled examples to make these measurements. However, these models often fail when faced with real-world imperfections because they have only seen "perfect" data during training. To solve this, the researchers turned to a technique called self-supervised learning. Instead of teaching the computer by showing it the correct answers, they taught it to recognize the underlying structure of the data by exposing it to thousands of corrupted versions of the same event. The computer learned to ignore the noise and focus on the core pattern, effectively becoming an expert at seeing through the distortion.

The researchers tested their method on a simulated dataset representing a heavy particle with a mass ranging from 2.5 to 6.5 trillion electron volts, a scale far heavier than any particle currently known. This particle was imagined to decay in a cascade, breaking down into a gluon, eight light quarks, a charged lepton, and a neutrino, which escapes the detector and is recorded only as missing energy. To prepare their AI, the team first fed it millions of these collision events without telling it what the mass of the parent particle was. During this phase, they deliberately introduced realistic errors into the data, mimicking the kinds of mistakes real detectors make. They randomly removed some particles, shifted the measured energy of others, blurred the angles, and even swapped the identities of particles to see if the AI could still recognize the event as a whole. The goal was to force the AI to build a mental map of these events that remained stable regardless of the corruption.

Once the AI had learned this robust, corruption-resistant way of seeing the data, the researchers then fine-tuned it with a small amount of labeled data to perform the actual task of mass regression. They compared this new approach against a traditional model trained from scratch on the same corrupted data. The results showed a clear advantage for the self-supervised method. While both models could predict the average mass of the particle, the new model produced a much sharper and more precise measurement. In the world of particle physics, a sharper measurement means a narrower peak in the data, which makes it significantly easier to distinguish a real discovery from background noise. The new model maintained this precision even when the data was heavily distorted, whereas the traditional model's performance degraded as the errors increased.

The study demonstrated that by learning to ignore the noise before learning the answer, the AI could reconstruct the mass of the heavy resonance with a resolution that was consistently better across all tested scenarios. When the researchers applied a standard correction to account for a slight bias in the average prediction, the new model still outperformed the traditional one, narrowing the width of the reconstructed mass by between 11% and 36% depending on the type of error present. This improvement is critical because the sensitivity of a search for new physics depends directly on how clearly the signal stands out. A method that can deliver a cleaner signal despite imperfect detectors offers a powerful new tool for future experiments. The researchers note that while their work was conducted on simulated data, the approach suggests a viable path forward for handling the complex, messy reality of high-energy physics, potentially allowing scientists to find new particles that were previously hidden in the blur of detector imperfections.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →