← Latest papers
⚛️ phenomenology

Prompting Particle Physics: Tokenized Multi-modal Foundation Models for Combinatorially Many Tasks

This paper introduces a tokenized multi-modal foundation model that unifies various collider reconstruction and simulation tasks into a single framework, enabling the generation of diverse intermediate physics objects through flexible prompt-based modality mapping while achieving state-of-the-art performance and Geant4-level fidelity.

Original authors: Nilotpal Kakati, Daniel Murnane, Baran Hashemi, Samuel Klein, Jeffrey Krupa, Eilam Gross, Lukas Heinrich, Michael Kagan

Published 2026-09-29
📖 5 min read🧠 Deep dive

Original authors: Nilotpal Kakati, Daniel Murnane, Baran Hashemi, Samuel Klein, Jeffrey Krupa, Eilam Gross, Lukas Heinrich, Michael Kagan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the heart of the Large Hadron Collider, a machine buried deep beneath the Swiss-French border, scientists smash protons together at nearly the speed of light. When these particles collide, they shatter into a shower of new, often fleeting, subatomic particles. To understand the fundamental laws of the universe, physicists must reconstruct what happened in that split second. They start with raw data: tiny electrical signals from sensors and energy deposits in massive detectors. From these scattered clues, they must build a complete picture of the collision, identifying the paths of charged particles, grouping energy into clusters, and finally piecing together the original particles that flew out of the crash. For decades, this reconstruction has been a long, winding chain of specialized computer programs. Each program in the chain is a master of a single task, tuned by hand to handle one specific step, from tracking a particle's path to measuring its energy. While effective, this approach is rigid, slow, and often loses information as data passes from one specialized tool to the next.

A team of researchers has now proposed a different way to look at this problem, treating the entire reconstruction process not as a chain of separate tools, but as a single, flexible conversation. They asked whether one computer model could learn to perform many of these steps at once, moving fluidly between different types of data. Instead of building a new algorithm for every job, they taught a single artificial intelligence to speak a universal language of particle physics. By converting every piece of information—whether it is a raw sensor reading, a particle's path, or a high-level energy measurement—into a simple list of discrete symbols, they created a shared vocabulary. In this new system, the model can be asked to take raw sensor data and produce a list of particles, or take a list of particles and simulate what the sensors would have seen. The model does not just guess the final answer; it generates the intermediate steps, producing tracks and clusters that physicists can inspect and verify, just as they do with traditional methods.

The researchers trained two versions of this model on a massive dataset of simulated particle collisions. One version, which they call nanoHEP, works like a storyteller, generating the output one symbol at a time, using the previous symbol to decide the next. The other, HEP4M, works more like a painter, predicting all the symbols for a given object simultaneously in a single glance. They tested these models on a staggering variety of tasks. In total, they explored over 1,300 different combinations of inputs and outputs. For example, they asked the model to take tracks and energy clusters and predict the true particles (a task known as particle flow), or to take true particles and predict the raw sensor signals (a task known as detector simulation). They even tested combinations the model had never seen before, such as feeding it four different types of data at once to see if it could combine them to make a better prediction.

The results show that this unified approach is not only possible but highly effective. The autoregressive model, which generates data step-by-step, proved particularly impressive. When tasked with reconstructing particles from detector data, it outperformed the current state-of-the-art specialized algorithm in reconstructing the jet momentum in terms of resolution and classifier score, while the single-task HEP4M model achieved the best calibrated response with a median closest to unity. More importantly, the objects it generated—such as the paths of charged particles and the shapes of energy clusters—were so realistic that they were much less separable from the data produced by the most detailed physics simulations available, though they remained distinguishable from the truth. The model learned the complex relationships between different types of data so well that it could generalize to new tasks. When presented with a combination of four input types it had never seen during training, it produced results that were at least as accurate as those achieved on every trained combination, matching or improving upon the performance of models trained specifically for those single tasks. This suggests that the model learned a deep, underlying understanding of how particle physics works, rather than just memorizing specific patterns.

However, the study also highlights a trade-off. The model that generates data step-by-step is more accurate but takes longer to run, while the model that predicts everything at once is faster but slightly less precise in capturing the fine details of the data. The researchers found that the best approach depends on the specific needs of the experiment. For tasks requiring the highest possible fidelity to the true physics, the slower, step-by-step model is superior. For tasks where speed is critical, the faster model offers a compelling alternative. Crucially, the researchers demonstrated that this single model can handle the entire spectrum of reconstruction and simulation tasks, from the rawest sensor data to the highest-level physics features, without needing to be retrained for each new job.

This work represents a significant shift in how particle physics data is processed. By moving away from a collection of specialized, isolated tools toward a single, multi-purpose foundation model, scientists can potentially streamline the entire analysis pipeline. The model's ability to produce intermediate objects that are physically interpretable means it does not have to be a "black box"; physicists can still look at the tracks and clusters it generates to ensure they make sense. While the model still faces challenges, particularly in generating raw sensor data with perfect precision, the success of this approach suggests a future where a single, versatile intelligence can help unravel the complex stories hidden within the collisions of the subatomic world. The researchers have made their code and data available, inviting the broader scientific community to build upon this new way of seeing the universe.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →