← Latest papers
⚛️ high-energy experiments

Scaling Collider Event Generation with Residual-Quantized Tokens

This paper introduces a scalable, autoregressive transformer-based framework for generating full collider events using residual-quantized tokens, demonstrating that token-level loss effectively predicts physical fidelity and enabling fast, ML-based surrogates to overcome bottlenecks in High-Luminosity LHC simulations.

Original authors: Dan Godi, Dmitrii Kobylianskii, Eilam Gross

Published 2026-10-02
📖 5 min read🧠 Deep dive

Original authors: Dan Godi, Dmitrii Kobylianskii, Eilam Gross

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the heart of Europe, where the Large Hadron Collider smashes protons together at nearly the speed of light, scientists are facing a problem of sheer volume. When the machine reaches its full power in the coming years, it will generate so much data that the computers needed to simulate what happens inside the detector will run out of time and memory. To understand these collisions, researchers traditionally rely on massive, step-by-step computer programs that mimic the behavior of every single particle as it flies through the detector. These programs are incredibly accurate but also incredibly slow, requiring vast amounts of computing power just to create the synthetic data needed to test new theories. The scientific community needs a faster way to generate these simulations, one that can keep up with the machine without sacrificing the physical reality of the results.

A team of researchers at the Weizmann Institute of Science in Israel has proposed a solution by borrowing a technique from the world of language models. Just as modern artificial intelligence can write coherent stories by predicting the next word in a sentence, these scientists have trained a similar system to predict the next "piece" of a particle collision. Instead of dealing with the messy, continuous numbers that describe a particle's speed and position, they first translate every particle into a short, discrete sequence of symbols, much like turning a sentence into a string of letters. They then use a powerful computer model to learn the patterns of these symbol sequences, allowing it to generate new, realistic collision events by simply predicting the next symbol in the chain. This approach transforms the complex physics of a particle detector into a language problem, where the computer learns the grammar of subatomic interactions.

The researchers tested this method on data from the CMS detector, a massive instrument that records the debris of proton collisions. They started by taking real collision events and breaking them down into their fundamental components: the particles that emerge from the collision and the signals those particles leave behind in the detector. To make this data manageable for their model, they used a specialized tool to compress the continuous physical properties of each particle—such as its momentum and direction—into a fixed set of digital codes. This process is like translating a high-resolution photograph into a series of pixel values that a computer can easily store and manipulate. By training their model on millions of these translated events, the system learned to generate new sequences of codes that, when converted back into physical numbers, looked and behaved exactly like real detector data.

What makes this work particularly significant is how the researchers verified that the generated data was not just statistically similar, but physically faithful. They did not simply check if the average speed of the particles matched; they examined the entire structure of the events, from individual particles to the complex sprays of debris known as jets. They found that the model could successfully recreate the subtle, random fluctuations that occur when particles interact with the detector material. Crucially, the system was able to generate events for physical processes it had never seen before, such as specific types of rare particle decays, by conditioning its predictions on the initial collision parameters. This suggests the model learned the underlying rules of the detector's response rather than just memorizing the training data, a vital requirement for a tool that must handle the unknown physics of future discoveries.

The team also investigated how the size of the model and the amount of data it saw affected its performance. They trained versions of the system ranging from ten million to over one billion parameters, testing them on datasets of varying sizes. They discovered a clear and predictable relationship: as the models grew larger and saw more data, the error in their predictions dropped in a consistent, mathematical way. Perhaps most importantly, they found that a simple measure of how well the model predicted the next symbol in the sequence was a reliable indicator of how accurate the final physical results would be. This means that scientists can use the model's internal error rate to gauge the quality of the generated physics without needing to run expensive, full-scale simulations to check the output.

By demonstrating that these discrete, language-based models can faithfully reproduce the complex output of a particle detector, the researchers have provided a new pathway for scaling up collider simulations. Their work shows that it is possible to bypass the computational bottlenecks that threaten to slow down future discoveries. The method relies on a straightforward principle: if you can translate the language of physics into a format that a computer can read as a sequence of symbols, you can use the most powerful tools of artificial intelligence to speak that language back to you. The results suggest that in the era of the High-Luminosity Large Hadron Collider, the future of simulation may not be about building bigger computers to crunch numbers, but about teaching machines to understand the grammar of the universe.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →