"Transforming" LHCb: self-supervised maps of heavy-flavour decays
This paper proposes and validates a self-supervised transformer architecture that learns heavy-flavour decay maps from LHCb data without explicit labels, demonstrating improved tagging performance and the ability to detect decays with invisible particles by inferring masked particle information and identifying anomalies in reconstructed jets.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The universe is built from a handful of fundamental particles, but the most interesting stories often happen when these particles decay, transforming into lighter, more stable forms. In the high-energy collisions at the Large Hadron Collider, heavy versions of these particles, known as beauty and charm hadrons, are created in vast numbers. When they break apart, they leave behind a trail of other particles that detectors can see. Physicists study these trails to find cracks in our current understanding of nature, looking for signs of invisible particles or strange behaviors that the standard model of physics cannot explain. However, finding these subtle signals is difficult because the detectors only see the visible pieces of the puzzle. If a heavy particle decays into something that leaves no trace, or if the decay is messy and incomplete, the standard tools for identifying what happened often fail. The challenge is to learn how to read the full story of a decay just by looking at the fragments that remain, even when some of the story is missing.
Researchers at Brown University have developed a new way to teach computers how to understand these heavy-particle decays without needing a textbook of answers. Instead of showing the computer millions of examples with the correct labels, they let the machine learn the rules of the game by trying to fill in the blanks. They took a massive dataset of simulated collisions from the LHCb experiment and presented the computer with jets—sprays of particles created in the collision—where they deliberately removed some of the tracks or hid their identities. The computer's task was to guess what was missing. By repeatedly trying to reconstruct the full picture from incomplete information, the system learned a deep, internal map of how heavy particles usually break apart. This map captures the complex relationships between the particles, such as how they move, where they come from, and what they are, without ever being told the specific name of the parent particle.
Once the computer had learned this map, the researchers tested whether it actually understood the physics. They found that the system could distinguish between different types of heavy particles and even tell the difference between matter and antimatter versions of the same particle, performing just as well as traditional methods that rely on fully labeled data. More importantly, they tested if the system could spot when a real decay was incomplete. When they took a reconstructed decay from a real collision and removed a known piece of it, the computer's "missingness score" went up significantly, indicating that the jet looked less complete than it should. This reaction was much stronger than when they removed random, unrelated particles from the same jet. This suggests the system has learned the specific, coherent structure of a heavy-particle decay and can sense when that structure is broken.
The team then took this trained system and applied it to real data from proton-proton collisions recorded in 2017. Without any further training on this specific data, the system successfully identified the same patterns in real-world collisions. It could separate different types of heavy-flavor decays and responded to the removal of specific particles in the real data just as it did in the simulations. The researchers also used the system to explore the data for unusual patterns. They found a small group of jets that looked very different from the rest: they were packed with charged particles but had almost no neutral ones. While the researchers could not immediately say what caused this, the system successfully organized these strange events together, showing that the method can group similar anomalies for further investigation.
This work demonstrates that a computer can learn the language of particle decay directly from the data itself, without needing a human to write the dictionary first. By learning to complete the picture of a decay, the system gains a powerful ability to recognize when something is wrong or missing. This approach offers a new way to search for physics beyond our current understanding, particularly for rare or incomplete decays that might otherwise be hidden in the noise. The researchers have shown that by teaching machines to understand the structure of these events, we can open new windows into the invisible parts of the universe, potentially revealing new particles or forces that have so far remained out of reach.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.