SoLiD26: A First Principles Solid-Liquid Interface Dataset for Machine-learned Interatomic Potentials
The paper introduces SoLiD26, a comprehensive dataset of 15.4 million first-principles atomic structures covering diverse solid-liquid interfaces, designed to train and benchmark machine-learned interatomic potentials for applications in electrochemistry, catalysis, and corrosion.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the hidden world of materials science, some of the most critical processes happen not in the center of a solid block or floating freely in a liquid, but right where the two meet. This boundary, known as the solid-liquid interface, is where batteries charge, where catalysts speed up chemical reactions, and where metals begin to corrode. To understand these processes, scientists rely on computer simulations that act as virtual microscopes, allowing them to watch atoms move and interact. For decades, the most accurate way to do this has been to solve complex equations based on the laws of quantum mechanics for every single atom in the system. While incredibly precise, this method is so computationally heavy that it can only simulate a tiny handful of atoms for a very short time, like watching a single frame of a movie. To see the full story of how a battery degrades or a catalyst works, researchers need to simulate millions of atoms over much longer periods, a task that traditional methods simply cannot handle.
To bridge this gap, scientists have developed a new generation of tools called machine-learned interatomic potentials. Think of these as smart shortcuts: a computer program is first taught by the slow, accurate quantum mechanics method, and then it learns to predict how atoms will behave with incredible speed, allowing for simulations of massive systems. However, for these shortcuts to work well in the messy, complex world where solids meet liquids, they need to be trained on data that specifically captures that unique environment. Until now, most training data has focused on isolated molecules or perfect crystals, leaving a gap in knowledge about how solids and liquids interact. A team of researchers at the Technical University of Denmark has now filled this gap by creating a massive, specialized library of data designed specifically for these interfaces.
The researchers, led by Jonas Busk and Tejs Vegge, compiled a dataset they call SoLiD26. This is not a small collection of examples; it is a vast archive containing over 15 million distinct atomic structures. Each structure represents a snapshot of a system where solid metal meets liquid water or electrolyte, capturing the chaotic dance of atoms as they arrange themselves, bond, and break apart. The dataset includes systems with up to 576 atoms and features 15 different chemical elements, ranging from common hydrogen and oxygen to precious metals like gold, platinum, and copper. These snapshots were generated using a rigorous method called density functional theory, which calculates the energy and forces acting on every atom with high precision. The team gathered these calculations from numerous previous studies on aqueous metal interfaces and electrochemical systems, then carefully cleaned and organized them into a single, coherent resource.
The process of building this dataset was meticulous. The researchers started by collecting raw data from various projects, but they knew that not all data is created equal. They filtered out calculations that used inconsistent settings or contained errors, such as structures where the forces on atoms were unrealistically high. They also removed duplicate entries to ensure the machine learning models would not be biased by seeing the same example too many times. To catch subtle errors that might slip past simple checks, they used a preliminary machine learning model to scan the data for outliers—strange structures that the model could not predict well, which often indicated a problem with the original calculation. This rigorous cleaning process ensured that the final library contained only high-quality, reliable information, making it a trustworthy foundation for training new AI models.
The true test of this dataset was to see if it could actually teach a machine learning model to understand solid-liquid interfaces better than existing models. The team took a powerful, pre-existing AI model and trained it on their new data. They compared this new model against the original pre-trained version and a model trained from scratch on the new data. The results were clear: the models trained on SoLiD26 made significantly fewer errors when predicting the energy and forces within these complex systems. Specifically, the error in predicting how atoms push and pull on each other dropped to about one-third of what it was before. This improvement suggests that the dataset successfully taught the AI the unique rules that govern how solids and liquids behave together, rules that were missing from the general training data used previously.
While the dataset is a significant step forward, the researchers are careful to note that it is a tool for development rather than a final solution. The models trained on this data still need further refinement, and the current tests focused on specific types of chemical elements and system sizes. However, the availability of this 100-gigabyte library, which includes detailed records of atomic positions, forces, and energies, opens the door for scientists worldwide to build better simulations. By providing a shared, high-quality resource for the solid-liquid interface, SoLiD26 allows researchers to move beyond the limitations of small-scale simulations and begin to model the complex, real-world materials that power our energy future. The work demonstrates that with the right data, machine learning can finally tackle the intricate boundaries where chemistry and physics collide.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.