Performance Portable Lattice Gauge Theory Simulation with Kokkos
This paper introduces \texttt{kwqft}, a performance-portable Kokkos-based implementation of Wilson pure gauge Monte Carlo simulations for Yang-Mills theory that supports arbitrary dimensions and multiple hardware backends while achieving accuracy and competitive performance across diverse high-performance computing systems.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
To understand the universe at its most fundamental level, physicists often turn to a method that treats space and time not as a smooth, continuous fabric, but as a grid of tiny, discrete points. Imagine a vast, invisible chessboard stretching across the cosmos, where every intersection holds a piece of information about the forces that bind matter together. This is the essence of lattice gauge theory, a powerful tool used to simulate the behavior of subatomic particles and the forces that govern them, particularly the strong force that holds atomic nuclei together. By breaking the universe down into these manageable chunks, scientists can use massive computers to trace the paths of particles and calculate how they interact, effectively running the laws of physics in reverse to see what emerges. However, the hardware used to run these simulations has become incredibly diverse. Today's supercomputers are a patchwork of different processors, from standard computer chips to specialized graphics cards, each speaking a different technical language and requiring its own unique set of instructions to run efficiently.
This diversity creates a significant headache for researchers. To get the best performance out of a specific machine, they often have to rewrite their simulation code from scratch for every different type of processor they want to use. Maintaining separate versions of the same algorithm for different computers is slow, expensive, and prone to errors. The question facing the scientific community is whether it is possible to write a single version of this complex simulation code that can run efficiently on any modern computer, from a single processor to a massive supercomputer, without sacrificing speed. Wei Sun, a researcher at the Institute of High Energy Physics in Beijing, has answered this question with a new software tool called kwqft. This program successfully demonstrates that it is possible to create a single, universal code for simulating the strong force that performs just as well as specialized, machine-specific code, regardless of the hardware it is running on.
The core of this achievement lies in a clever approach to software design that treats the computer's architecture as a variable rather than a fixed constraint. The researchers built their simulation using a modern programming framework that acts as a universal translator. Instead of writing separate instructions for a graphics card, a standard processor, or a specialized chip, they wrote one set of instructions that the framework automatically adapts to the machine it is running on. This allows the same code to run on a wide variety of systems, including those from different manufacturers, without the researchers needing to touch the source code again. The software is designed to be flexible, allowing scientists to change the number of dimensions in their simulation or the specific properties of the particles they are studying just by adjusting a few settings before the program starts, rather than rewriting the entire program.
To prove that this universal approach works, the team subjected their software to rigorous testing across a wide range of hardware. They ran simulations on a powerful graphics card from NVIDIA, a specialized chip from Hygon, and several different types of central processors, including advanced models from Huawei and Intel. In every case, the software produced results that matched the known, exact solutions for simpler versions of the problem and agreed with previously published data for more complex scenarios. The researchers verified that the simulation correctly reproduced the behavior of the strong force in two, three, and four dimensions, and for different types of particle groups, confirming that the universal code was not losing accuracy in the process of becoming portable.
The performance results were particularly striking. When running on a high-end graphics card, the universal software achieved nearly the same speed as a version of the code that had been hand-tuned specifically for that card, reaching over 90 percent of the specialized code's efficiency on larger simulations. On a massive supercomputer in Shenzhen, known as LineShine, the software scaled up to run across 512 nodes, or individual computing units, simultaneously. As the researchers added more computing power, the simulation sped up significantly, demonstrating that the software could handle the complex communication required to keep thousands of processors working together in harmony. Furthermore, on the latest generation of processors designed for high-speed vector calculations, the software was able to process multiple points on the grid at once, speeding up the simulation by a factor of three compared to a standard version of the same code.
The implications of this work extend beyond a single successful test. By proving that a single codebase can deliver competitive performance across such a diverse landscape of hardware, the researchers have removed a major barrier to scientific progress. Scientists no longer need to choose between writing code that is easy to maintain and code that is fast; they can now write once and run anywhere, confident that the software will adapt to the machine it is on. This opens the door for more complex simulations of the universe, allowing researchers to explore questions about the early moments of the cosmos and the nature of matter with greater ease and on a larger scale than ever before. The success of kwqft suggests that the future of high-performance computing in physics may not lie in constantly rewriting code for new machines, but in creating flexible, intelligent tools that can navigate the changing landscape of technology on their own.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.