← Latest papers
🔬 materials science

Small-supercell and Small-dataset Training Strategy of Machine Learning Interatomic Potentials for Point Defects

This paper proposes a computationally efficient machine learning interatomic potential training strategy that utilizes a small dataset derived from limited DFT calculations on small supercells, combined with multi-size defect and pristine bulk data, to accurately predict point defect properties in large-scale systems with minimal error.

Original authors: Zhenxing Dai, Mingjue Ni, Xinpeng Li, Menglin Huang, Anderson Janotti, Shiyou Chen

Published 2026-09-22
📖 5 min read🧠 Deep dive

Original authors: Zhenxing Dai, Mingjue Ni, Xinpeng Li, Menglin Huang, Anderson Janotti, Shiyou Chen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Inside the solid materials that power our modern world, from the silicon chips in our phones to the solar cells on our roofs, tiny imperfections dictate performance. These imperfections, known as point defects, are essentially missing atoms or misplaced ones within a crystal's orderly grid. While they are microscopic, their influence is macroscopic; they determine how well a semiconductor conducts electricity or how efficiently it converts light into energy. For decades, scientists have relied on complex computer simulations based on the laws of quantum mechanics to understand these defects. These simulations act like a digital microscope, allowing researchers to see how atoms shift and how energy changes when a defect is present. However, this digital microscope is incredibly slow and expensive to use. To get accurate results, especially for defects that carry an electric charge, the simulations must model a vast number of atoms—often hundreds or even thousands—to avoid artificial errors caused by the boundaries of the computer model. This requirement makes it nearly impossible to study these defects in the large, realistic systems where they actually function.

A newer approach has emerged to speed up these calculations: machine learning interatomic potentials. Think of these as smart shortcuts. Instead of solving the heavy quantum equations every time, a computer program learns the rules of how atoms interact from a small set of high-quality examples. Once trained, this program can predict how atoms will behave in massive systems with near-perfect accuracy but in a fraction of the time. The challenge, however, has been that training these programs usually requires a massive library of data, often involving thousands of expensive simulations. This is particularly difficult for charged defects, where the need for huge atomic models to avoid errors makes gathering data prohibitively costly. Researchers at Fudan University and the University of Delaware have now developed a way to train these machine learning models using a tiny fraction of the usual data, proving that a small, carefully chosen set of examples can teach the computer to predict the behavior of defects in much larger systems.

The team focused on a specific strategy to see if they could build a reliable model without the usual mountain of data. They chose three different materials—gallium nitride, silicon dioxide, and a copper-based compound used in solar cells—as their test subjects. For each, they looked at specific defects, such as a missing nitrogen atom or a misplaced copper atom, in various electric charge states. Their goal was to see if they could train a machine learning model using only a few small atomic models, containing fewer than one hundred atoms, and then have that model accurately predict what would happen in a much larger model, containing over two hundred atoms. They tested two different ways of feeding data to the computer. In the first method, they provided only data from the defective atoms themselves, taken from simulations of different small sizes. In the second method, they added data from perfect, defect-free chunks of the same material to help the computer understand the background environment.

The results revealed a clear lesson about how these digital shortcuts learn. When the researchers trained the model using data from just one small size of atomic model, the program failed miserably when asked to predict the behavior of larger systems. The errors in energy predictions were enormous, sometimes reaching tens of electronvolts, which is a massive mistake in the world of atomic physics. The computer had simply memorized the specific quirks of that small model rather than learning the underlying physics. However, when the training data included examples from several different small sizes, the model's ability to predict the larger systems improved dramatically. By seeing how the defect behaved in a small cluster, a medium cluster, and a slightly larger one, the computer learned to recognize the patterns that hold true regardless of the system's size. This allowed it to make accurate predictions for systems containing over two hundred atoms, with energy errors dropping to less than 0.3 electronvolts in many cases.

Adding perfect, defect-free material to the training data provided an extra boost, particularly for defects with a low electric charge. This extra information helped the model understand the stable, undisturbed environment that surrounds the defect, making its predictions even more reliable. Yet, the study also highlighted a limit to this efficiency. For defects with a very high electric charge, simply using small models and adding perfect material was not enough. The long-range electrical forces in these highly charged systems are so strong that the model still needed to see examples of larger, charged-defect systems to learn the correct behavior. Without these larger examples, the predictions remained unstable.

This work provides a practical roadmap for scientists who want to use machine learning to study defects without waiting years for supercomputers to finish their calculations. It shows that you do not need a massive dataset to get good results; you need a smart one. By carefully selecting a few examples from different sizes and mixing in data from perfect materials, researchers can train models that are accurate enough to simulate large-scale defects. This approach makes it feasible to study the microscopic flaws that control the performance of the electronic and optical devices that shape our daily lives, turning a computationally impossible task into a manageable one.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →