← Latest papers
🧬 biology

Breaking the Exascale Barrier for the Electronic Structure Problem in Ab-Initio Molecular Dynamics

This paper demonstrates that a modified non-orthogonal local submatrix method achieves over 1.1 EFLOP/s on 4,400 NVIDIA A100 GPUs, enabling ab-initio molecular dynamics simulations of SARS-CoV-2 spike proteins containing up to 83 million atoms.

Original authors: Robert Schade, Tobias Kenter, Hossam Elgabarty, Michael Lass, Thomas D. Kühne, Christian Plessl

Published 2026-09-29
📖 6 min read🧠 Deep dive

Original authors: Robert Schade, Tobias Kenter, Hossam Elgabarty, Michael Lass, Thomas D. Kühne, Christian Plessl

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

To understand how matter behaves at its most fundamental level, scientists often turn to a technique called ab-initio molecular dynamics. This approach attempts to simulate the movement of atoms in molecules, surfaces, or solids by calculating the forces that push and pull them. Unlike older methods that rely on simplified, empirical rules, this technique solves the complex quantum-mechanical problem of electrons directly. It treats the electrons not as a vague background, but as the primary actors whose behavior dictates how atoms interact. While this provides a highly accurate picture of chemical reactions and material properties, it comes with a steep price: the computational effort required grows so rapidly with the size of the system that simulating large, realistic structures has long been considered impossible. For decades, researchers were forced to study tiny fragments of matter, unable to see the full picture of how billions of atoms move together in a living cell or a complex material.

A team of researchers from Paderborn University in Germany has now pushed past this barrier, demonstrating a way to simulate electronic structures for systems containing up to 83 million atoms. By adapting a method known as the non-orthogonal local submatrix technique and running it on a massive supercomputer equipped with thousands of specialized graphics processors, they achieved a sustained computing speed of over 1.1 exaflops. This means the system performed more than one quintillion floating-point operations per second, a milestone that places this specific scientific application among the very first to break the "exascale" barrier. The researchers did not merely run a simulation; they engineered a new way to organize the mathematical work so that the hardware could operate at nearly 80 percent of its maximum theoretical capacity. Their work proves that with the right algorithmic adjustments, it is possible to calculate the quantum behavior of entire proteins in solution, opening the door to studying biological machinery and complex materials with unprecedented detail.

The core challenge in these simulations is that every time an atom moves, even by a tiny fraction, the entire electronic structure of the system must be recalculated to determine the new forces acting on that atom. In traditional approaches, this recalculation becomes prohibitively slow as the number of atoms increases. The researchers utilized a method that breaks the massive mathematical problem into smaller, manageable pieces called submatrices. Instead of trying to solve the equation for the entire system at once, the computer isolates small sections of the data, solves them independently, and then reassembles the results. This "local" approach avoids the need for constant communication between different parts of the computer, which is usually the bottleneck in large-scale simulations. The team applied this method to a specific biological target: the spike protein of the SARS-CoV-2 virus, which is the structure the virus uses to attach to human cells. They simulated the protein anchored in a lipid layer and surrounded by water, creating a system with approximately 1.7 million atoms for their initial tests, and then scaling up to a grid of these proteins to reach a total of 83 million atoms.

To achieve this level of performance, the researchers had to go beyond simply using more computers; they had to fundamentally rethink how the software interacts with the hardware. The supercomputer they used, located at the National Energy Research Scientific Computing Center, is equipped with 4,400 NVIDIA A100 graphics processing units. These chips are designed to handle massive amounts of parallel calculations, but they are most efficient when working on large, dense blocks of data. The original version of the algorithm created submatrices that were often too small to fully utilize the power of these chips. The team introduced a new heuristic, or a set of rules for decision-making, that combined multiple columns of data into larger submatrices before processing them. This adjustment was not just a minor tweak; it was a strategic shift based on the specific performance characteristics of the graphics processors. By measuring how fast the chips could multiply matrices of different sizes, the researchers optimized the grouping of data to ensure the hardware was working at its peak efficiency. This modification allowed the system to sustain a performance level that was previously thought unattainable for this type of calculation.

The results of this effort were measured by running the simulation on a grid of spike proteins, effectively creating a virtual environment with 83 million atoms. The team tracked the time it took to complete each step of the calculation and the total number of mathematical operations performed. They found that the system consistently delivered a speed of between 1.106 and 1.127 exaflops, maintaining about 80 percent of the theoretical peak performance of the hardware. This is a significant achievement because high-performance computing systems often struggle to maintain high efficiency when scaling up to thousands of processors; usually, the efficiency drops as communication overhead increases. In this case, the local nature of the method meant that the processors spent almost all their time calculating rather than waiting for data. The researchers verified that the accuracy of the simulation remained high, ensuring that the speed gains did not come at the cost of scientific validity.

This breakthrough is not just about simulating a virus; it represents a new capability for computational science. The ability to model electronic structures for systems with tens of millions of atoms means that scientists can now study phenomena that were previously out of reach, such as the behavior of large biomolecules in their natural, watery environments or the properties of complex materials under stress. The method is versatile enough to be applied to any problem that requires evaluating a mathematical function for a large, sparse set of data, not just molecular dynamics. By demonstrating that exascale performance is achievable in a real-world scientific application, the researchers have provided a blueprint for future high-performance computing. They have shown that by aligning algorithmic design with hardware capabilities, it is possible to solve problems that were once considered too large to compute, bringing the quantum mechanical world of atoms and electrons into clearer focus for the study of life and matter.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →