Surrogate model driven RL optimization of the TWOCRYST crystal angular alignment
This paper introduces MDHaloEnv, a Gymnasium-based reinforcement learning environment utilizing Proximal Policy Optimization and theoretically grounded Markov Decision Process improvements to automate and enhance the microradian-level alignment of crystals in the TWOCRYST experiment, thereby improving commissioning efficiency for future LHC-based double-channeling studies.