← Latest papers
🔬 physics

Surrogate model driven RL optimization of the TWOCRYST crystal angular alignment

This paper introduces MDHaloEnv, a Gymnasium-based reinforcement learning environment utilizing Proximal Policy Optimization and theoretically grounded Markov Decision Process improvements to automate and enhance the microradian-level alignment of crystals in the TWOCRYST experiment, thereby improving commissioning efficiency for future LHC-based double-channeling studies.

Original authors: Leander Grech, Gianluca Valentino, Oliver Aberle, Federico Alessio, Chiara Antuono, Gianluigi Arduini, Laura Bandiera, Massimo Benettoni, Marco Calviani, Sara Cesare, Victor Coco, Simone Coelli, Georg
Published 2026-09-01
📖 6 min read🧠 Deep dive

Original authors: Leander Grech, Gianluca Valentino, Oliver Aberle, Federico Alessio, Chiara Antuono, Gianluigi Arduini, Laura Bandiera, Massimo Benettoni, Marco Calviani, Sara Cesare, Victor Coco, Simone Coelli, Georges Daher, Quentin Demassieux, Davide De Salvador, Mario Di Castro, Luigi Esposito, Massimiliano Ferro-Luzzi, Alex Fomin, Jianling Fu, Paolo Gandini, Vincenzo Guidi, Hana Havlikova, Pascal Hermes, Sergio Jaimes Elles, Sune Jakobsen, Krzysztof Korcyl, Gianluca Lamanna, Simone Libralon, Chiara Maccani, Lorenzo Malagutti, Daniele Marangotto, Fernando Martinez Vidal, Eloise Matheson, Jose Mazorra de Cos, Andrea Mazzolari, Andrea Merli, Haixing Miao, Daniele Mirarchi, Riccardo Negrello, Nicola Neri, Marcin Patecki, Antonio Perillo Marcone, Jacopo Pinzino, Stefano Redaelli, Patrick Robbe, Marco Romagnoni, Benoit Salvant, Izaac Sanderswood, Philippe Schoofs, Regis Seidenbinder, Gabriele Simi, Santiago Solis Paiva, Enrica Soria, Marco Sozzi, Elisabetta Spadaro Norella, Achille Stocchi, Katie Taylor, Giorgia Tonani, Andrea Triossi, Nicola Turini, Santiago Vico Gil, Chiara Vinotto, Volodymyr Svintozelskyi, Tianyu Xing, Marco Zanetti, Federico Zangari, Carlo Zannini

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Inside the massive, ring-shaped tunnel of the Large Hadron Collider, protons travel at nearly the speed of light, circling billions of times per second. To study the fundamental building blocks of the universe, scientists often need to peel away the outer layers of this beam, extracting a thin halo of stray particles and steering them toward a specific target. This is a delicate operation. The particles must be guided with extreme precision using bent silicon crystals, which act like microscopic mirrors that trap and bend the protons. However, these crystals must be aligned with an accuracy finer than the width of a human hair. If they are even slightly off, the beam misses its mark, and the experiment fails. Traditionally, aligning these crystals has been a slow, manual process, requiring expert operators to spend hours making tiny adjustments while watching for faint signals. It is a task that demands patience and a steady hand, but it is also a bottleneck that limits how quickly new experiments can begin.

A team of researchers from CERN and universities across Europe has now developed a new way to solve this problem. They created a computer program that learns to align these crystals automatically, using a method called reinforcement learning. Instead of relying on complex physics equations to predict how the beam will behave, the program learns directly from real data collected during previous tests. It treats the alignment process like a game: the computer acts as a player that makes small adjustments to the crystal's angle, and it receives points based on how well the beam hits the target. Over time, the program learns the best moves to maximize its score, eventually discovering a strategy that is faster and more consistent than human operators. This approach, tested on the TWOCRYST experiment, represents a significant step toward making particle accelerators more autonomous and efficient.

The core of this achievement is a virtual environment the researchers built to train their artificial intelligence. Because the real accelerator is too valuable and too risky to use for trial-and-error learning, the team created a digital twin called MDHaloEnv. This system takes historical data from beam tests, where crystals were manually rotated through thousands of positions, and uses it to simulate what would happen if the computer tried to align the crystal itself. The program receives images from a high-speed camera that tracks the particles, along with data from sensors that measure beam loss. It then decides whether to turn the crystal slightly left or right. If the adjustment brings the beam closer to the ideal path, the program gets a reward; if it moves away, it gets a penalty.

To make this learning process effective, the researchers had to overcome several challenges. The data they used was imperfect, containing gaps and inconsistencies because it came from real-world experiments rather than a perfect simulation. The signals from the sensors were often noisy, making it difficult to tell if a small change in the beam was due to a good adjustment or just random fluctuation. To handle this, the team designed a reward system that smooths out these fluctuations, giving the computer a clearer picture of whether it was moving in the right direction. They also programmed the system to avoid making unnecessary, jittery movements, encouraging it to find a stable position rather than oscillating back and forth.

When they trained their computer agent using this setup, the results were impressive. The agent learned to navigate the complex landscape of crystal angles and consistently found the sweet spot where the beam was perfectly aligned. In tests, the program could start from a wide range of initial positions—some far away from the target—and successfully guide the crystal to the correct alignment within a few minutes. It learned to recognize the subtle patterns in the camera images that indicated the beam was in the right place, even when the signals were weak or noisy. The system proved robust, maintaining its performance even when the starting conditions changed, suggesting it had learned a general strategy rather than just memorizing a specific path.

The researchers compared their new method against different variations of the reward system to see which worked best. They found that combining a smoothing technique with a penalty for excessive movement produced the most stable results. The computer learned to settle into the optimal position with very little shaking, reaching a level of precision that matched the theoretical limits of the data itself. This means the system was not just guessing; it had truly learned the underlying behavior of the beam and the crystal. The success of this approach suggests that similar techniques could be applied to other parts of the accelerator, potentially allowing for more complex experiments that require multiple crystals to be aligned simultaneously.

This work is particularly important for future experiments, such as the proposed ALADDIN experiment, which aims to measure the magnetic and electric properties of short-lived particles. These experiments will rely on the same double-crystal setup and will require even faster and more precise alignment than what is currently possible. By proving that a computer can learn to do this task from real data without needing a perfect physics model, the researchers have opened the door to a new era of autonomous control in particle physics. The system does not need to understand the deep mathematics of particle interactions; it simply needs to know what a successful alignment looks like, and it can learn that from experience.

The path forward involves testing this system in real-time during future beam tests. The team plans to run the computer in a "shadow mode," where it observes the experiment and suggests adjustments without actually controlling the hardware. This will allow them to compare the computer's decisions with those of human operators and ensure safety before handing over full control. If these tests are successful, the method could become a standard tool for commissioning new experiments, saving valuable time and allowing scientists to focus on the physics rather than the mechanics of alignment. The ability to automate such a critical and delicate task marks a significant evolution in how we operate the world's most powerful scientific instruments.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →