← Latest papers
🔬 materials science

The Roadmap of Inorganic Computational Materials Databases: Capabilities, Credibility, Coverage, and the Open Frontier

This paper evaluates the current landscape of inorganic computational materials databases, identifying that while methodological capabilities are mature, the primary barrier to expansion is the economic cost of achieving trustworthy accuracy, and proposes a three-horizon roadmap to systematically address these gaps and unlock high-cost property families through advanced workflows and community governance.

Original authors: Miao Liu, Jianghao Jin, Tenglong Lu, Jianguo Si, Yin Shi, Sheng Meng, Weihua Wang

Published 2026-09-18
📖 5 min read🧠 Deep dive

Original authors: Miao Liu, Jianghao Jin, Tenglong Lu, Jianguo Si, Yin Shi, Sheng Meng, Weihua Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

For the past fifteen years, the search for new inorganic materials has undergone a quiet revolution. Instead of testing one substance at a time in a laboratory, scientists now use powerful computers to calculate the properties of entire families of chemical compounds before a single atom is ever mixed. This approach relies on a method called density functional theory, a set of mathematical rules that allows computers to predict how electrons behave inside a solid. By running these calculations, researchers can generate vast libraries of data, effectively creating a digital map of the material world. These databases have become essential tools, guiding experimentalists toward promising candidates for batteries, solar cells, and superconductors without them needing to run a single simulation themselves. The promise is that by mapping out these possibilities, we can discover useful materials faster and more efficiently than ever before.

However, a new perspective from a team of researchers at the Institute of Physics in China reveals that this digital map is far from complete. While the field has grown rapidly, the growth has been uneven, leaving large territories of the material world uncharted. The authors surveyed nineteen different types of material properties, ranging from simple measurements like how a crystal is shaped to complex behaviors like how it conducts heat or responds to magnetic fields. They found that the technology to calculate these properties already exists; the software tools are mature and capable of handling nearly every property an experimentalist could measure. The barrier is no longer a lack of capability, but rather the economics of trust. Some properties are cheap and easy to calculate with high confidence, while others are so expensive to compute or so sensitive to the details of the calculation that publishing them at a massive scale remains a gamble.

The researchers organized these properties into a landscape based on cost and reliability. At the top of this landscape, where calculations are both affordable and trustworthy, sit the properties that have already been harvested in large numbers. These include the basic shape of a crystal, its formation energy, and its stiffness. Major databases now contain millions of entries for these specific traits. Moving down the landscape, the picture changes. Properties like the precise energy needed to jump an electron across a gap, or the way a material vibrates to conduct heat, are harder to calculate. They require more computing power and more careful handling, so they appear in databases only for small, curated groups of materials. At the bottom of the landscape lie nine families of properties that are currently blank zones. These include nuclear magnetic resonance parameters, which tell us about the local magnetic environment of atoms, and quantum transport properties, which describe how electricity flows through tiny devices. For these, there are no systematic databases at all.

The authors argue that these empty spaces are not accidents of history but the result of rational decisions about what is worth the effort. Calculating a property like thermal conductivity, for instance, can take hundreds of thousands of hours of computer time for a single material, making it impossible to run for millions of compounds with current methods. Similarly, some properties depend heavily on specific choices made by the scientist running the simulation, meaning that without a strict, standardized protocol, the data would be too unreliable to trust. The paper suggests that the next decade of progress will not come from inventing new ways to calculate these things, but from finding smarter ways to do it. The roadmap they propose involves three stages. First, the community should focus on connecting the existing databases and standardizing how data is shared, ensuring that a number from one database can be compared directly with a number from another.

In the medium term, the goal is to industrialize the properties that are currently too expensive to calculate for everyone. The authors point to the rise of machine-learned interatomic potentials as a game-changer. These are models trained on a small set of highly accurate calculations that can then predict the behavior of millions of other materials with much less computing power. One recent study already used this approach to calculate thermal conductivity for over 230,000 materials, a feat that would have been impossible just a few years ago. By using these surrogate models, researchers can fill in the middle ground of the landscape, bringing properties like full phonon spectra and defect energies into the realm of large-scale data.

Looking further ahead, the final frontier involves conquering the most difficult properties, such as those related to complex spectroscopy and quantum transport. This will require not just better models, but a complete overhaul of the infrastructure. The authors envision a future where autonomous computing systems run simulations, verify results against experimental data, and even extract structured information from decades of scientific literature to build a complementary database. This would turn the vast archive of published papers into a queryable resource, allowing computers to learn from human discoveries directly. The ultimate goal is to transform the field from a collection of isolated projects into a coherent, trustworthy infrastructure. The first act of this story is complete: the tools are ready, and the easy data has been collected. The second act, which is just beginning, will be defined by how the community tackles the difficult, expensive, and currently blank zones to build a truly complete map of the inorganic world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →