← Latest papers
🔬 physics

MGKDB: An IMAS-aligned multicode gyrokinetic simulation database for reproducible fusion turbulence modeling and data-driven analysis

The paper introduces MGKDB, an open-source framework and curated archive that converts heterogeneous fusion simulation data into IMAS-aligned, traceable scientific records to enable reproducible, large-scale analysis, cross-code comparison, and data-driven modeling across multiple gyrokinetic codes.

Original authors: Craig Michoski, David R. Hatch, Dongyang Kuang, Matthew Waller, Chris Holland, M. J. Pueschel, Joseph McClenaghan, Tom F. Neiser, Max T. Curie, Venkitesh Ayyar, Joseph Schmidt, Leonhard A. Leppin, Aar
Published 2026-09-04
📖 6 min read🧠 Deep dive

Original authors: Craig Michoski, David R. Hatch, Dongyang Kuang, Matthew Waller, Chris Holland, M. J. Pueschel, Joseph McClenaghan, Tom F. Neiser, Max T. Curie, Venkitesh Ayyar, Joseph Schmidt, Leonhard A. Leppin, Aaron Ho, Nathan T. Howard, Tapan Ganatma Nakkina, Bhavin Patel, Yann Camenen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Inside the heart of a fusion reactor, a superheated gas called plasma swirls in a magnetic cage, holding the promise of limitless clean energy. To make this promise a reality, scientists must understand how tiny, chaotic whirlpools within that plasma steal heat away from the core, cooling the reaction before it can sustain itself. To predict these turbulent movements, researchers run incredibly complex computer simulations. These programs act as virtual laboratories, solving equations that describe how particles move and interact under extreme conditions. However, these simulations are expensive and time-consuming, often taking days or weeks of supercomputer time to complete a single run. For decades, the results of these massive calculations have been stored in isolated folders, tied to the specific software that created them. If a scientist wanted to compare a result from one program with another, or reuse an old calculation for a new study, they often faced a wall of incompatible file formats and missing details, forcing them to start from scratch.

A new system called MGKDB is changing how this data is handled, turning a scattered collection of isolated files into a unified, searchable library of scientific records. The researchers behind this project built a framework that takes the raw, complex outputs from different simulation programs and translates them into a common language without throwing away the original details. Imagine a massive archive where every entry keeps its original, detailed blueprint while also having a standardized summary card that anyone can read. This allows scientists to ask broad questions across thousands of simulations at once, such as finding all instances where a specific type of turbulence occurred, or checking how well different computer models agree with each other. The system preserves the full history of how each calculation was made, ensuring that the data remains trustworthy and reproducible for future work.

The core of this achievement is the ability to link three distinct types of information for every single simulation run. First, the system keeps the original, code-specific files that contain the exact instructions and raw numbers generated by the software. These are the authoritative records needed to reproduce the calculation exactly as it was done. Second, it attaches detailed metadata that tracks who ran the simulation, when, and with what specific settings, creating a clear chain of custody. Third, and perhaps most importantly, it converts the results into a standardized format based on a shared international framework known as IMAS. This translation layer allows scientists to search for physical quantities, like the rate of heat loss or the strength of magnetic fields, using the same terms regardless of which computer program generated the data. By keeping the original files alongside this common translation, the system ensures that scientists can discover patterns across different models without losing the specific context needed to interpret them correctly.

The researchers tested this new infrastructure by looking at over one million simulation records that had been gathered from three different types of fusion modeling tools. They found that nearly every single record included a complete, standardized version of the physics data, proving that the system could handle a massive volume of diverse inputs. To show what this library could actually do, they ran three specific demonstrations. In the first, they analyzed thousands of linear simulations to map out the different types of unstable waves that can form in the plasma. By looking at the standardized data, they were able to group these waves into distinct families based on their physical behavior, revealing that some types of turbulence are very similar to each other while others are quite different. This kind of large-scale pattern recognition would have been nearly impossible if the data remained locked in separate, incompatible formats.

In a second demonstration, the team examined the "geography" of the data to see where the simulations had been focused and where gaps remained. They mapped out the conditions under which the simulations were run, such as the temperature and density of the plasma, to see how well the archive covered the full range of possible scenarios. They discovered that the existing data was heavily clustered around certain conditions, reflecting the specific goals of past research campaigns, while other areas were left largely unexplored. This insight is crucial for planning future experiments, as it helps scientists identify which new simulations would provide the most valuable new information rather than simply repeating what is already known. The system also revealed that many simulations shared identical starting conditions, a detail that is vital for ensuring that statistical analyses do not accidentally count the same data point multiple times.

The final test showed how the archive could bridge the gap between high-fidelity, expensive simulations and faster, simplified models. The researchers took the input settings from fifty archived, complex simulations and used them to run a much faster, reduced-model calculation. Because the system preserved the exact link between the original complex run and the new simplified one, they could immediately compare the results. This process created a clean, ready-to-use dataset that could be used to train artificial intelligence models to predict plasma behavior. The key finding here was not just that the models could be compared, but that the entire workflow—from retrieving the old data to generating new results—could be automated and traced back to its source. This means that future scientists can build on these connections without having to manually reconstruct the steps taken by previous researchers.

The success of MGKDB lies in its refusal to choose between preserving the past and enabling the future. By keeping the original, detailed files intact while simultaneously creating a standardized, searchable layer on top, the system respects the complexity of the science while making it accessible. It acknowledges that different computer models make different assumptions and that a simple match in a database field does not guarantee that two results are physically identical. Instead, it provides the tools to make those distinctions clear and auditable. As the archive continues to grow, adding new types of simulations and more data, this infrastructure ensures that the computational investment of today remains a usable resource for the discoveries of tomorrow. The project demonstrates that with the right organization, the accumulated knowledge of fusion research can become a living, cumulative foundation rather than a collection of forgotten files.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →