← Latest papers
🔬 materials science

uMOF: A Universal Database, Benchmark, and Machine Learning Interatomic Potentials for Metal-Organic Frameworks

The paper introduces uMOF, a comprehensive resource comprising the largest DFT dataset for metal-organic frameworks (MOFs), a literature-mined experimental benchmark, and two universal machine learning interatomic potentials that significantly outperform existing models in predicting complex gas adsorption properties by leveraging diverse training data and high-level theory.

Original authors: Théo Jaffrelot Inizan (Materials Sciences Division, Lawrence Berkeley National Laboratory, Berkeley, CA, USA, Bakar Institute of Digital Materials for the Planet, Division of Computing, Data Science
Published 2026-08-31
📖 5 min read🧠 Deep dive

Original authors: Théo Jaffrelot Inizan (Materials Sciences Division, Lawrence Berkeley National Laboratory, Berkeley, CA, USA, Bakar Institute of Digital Materials for the Planet, Division of Computing, Data Science, and Society, University of California, Berkeley, CA, USA), Prathami Divakar Kamath (Materials Sciences Division, Lawrence Berkeley National Laboratory, Berkeley, CA, USA, Department of Materials Science & Engineering, University of California, Berkeley, CA, USA), Alin Marin Elena (Scientific Computing Department, Science and Technology Facilities Council, Daresbury Laboratory, UK), Kristin A. Persson (Materials Sciences Division, Lawrence Berkeley National Laboratory, Berkeley, CA, USA, Department of Materials Science & Engineering, University of California, Berkeley, CA, USA)

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where the materials that power our future—filters that clean the air, batteries that store solar energy, or catalysts that turn carbon dioxide into fuel—can be designed on a computer before a single gram is ever synthesized in a lab. This is the promise of computational materials science, a field that relies on simulating how atoms interact to predict how a material will behave. For decades, scientists have used a powerful but incredibly expensive method called density functional theory to calculate these interactions. It is like taking a high-resolution photograph of every single atom in a structure; the picture is accurate, but the camera is so slow and heavy that it can only capture a tiny, still image of a few hundred atoms. To study complex, porous materials that contain thousands of atoms and need to be watched moving over time, this method is simply too slow.

In recent years, a new generation of artificial intelligence tools has emerged to solve this speed problem. These tools, known as machine learning interatomic potentials, act as a smart shortcut. They are trained on the expensive, high-quality photographs to learn the rules of atomic behavior, allowing them to predict how atoms will move and interact at a fraction of the cost. However, while these AI tools have become masters at predicting the behavior of simple, dense crystals, they have struggled with a specific class of materials called metal-organic frameworks. These are sponge-like structures made of metal nodes connected by organic chains, creating vast internal spaces perfect for trapping gas molecules. The challenge has been that the AI models were trained on data that didn't quite match the unique chemistry of these sponges, leading to inaccurate predictions about how well they would hold onto gases like carbon dioxide or hydrogen.

A team of researchers has now bridged this gap with a comprehensive new resource called uMOF. This project is not just a single discovery but a complete toolkit designed to make the simulation of these complex materials reliable for the first time. The team began by generating the largest and most accurate set of training data ever created for metal-organic frameworks. They calculated the energy and forces for over 85,000 different configurations of nearly 20,000 unique frameworks, covering 79 different chemical elements. Crucially, they used a more advanced level of theory than previous studies, one that better accounts for the subtle, long-range forces that hold these sponges together and allow them to trap gas molecules. This dataset includes not just static snapshots of the materials, but also simulations of them moving and vibrating at different temperatures, providing the AI with a dynamic view of how the structures behave in the real world.

To ensure their new tools actually worked, the researchers built a rigorous testing ground. They used advanced language models to scan hundreds of scientific papers, extracting thousands of verified experimental measurements for properties like how much gas a material can hold or how it expands when heated. This created a massive benchmark of real-world data against which they could test their AI models. They then trained two new AI models on their high-quality dataset. The results were striking. When tested on standard, stable properties like how stiff a material is, the new models performed just as well as existing tools. But when the tests moved to the harder, more dynamic task of predicting gas adsorption—how much gas a sponge can soak up and how tightly it holds it—the new models left the competition far behind.

The new models reduced the error in predicting gas binding strength by more than 80 percent compared to other specialized models, bringing the predictions down to the level of experimental uncertainty. This success was not just about having more data; in fact, one of the competing models was trained on a dataset nearly a thousand times larger, yet it performed worse. The researchers found that the key to their success was the quality of the physics in their training data and the inclusion of a small but critical fraction of high-temperature movement simulations. Even though these moving snapshots made up only about 1.7 percent of their total data, they were essential for teaching the AI how to handle the chaotic, high-energy moments that occur when gas molecules collide with the framework. Without this specific type of data, the models failed to capture the stability required for real-world applications.

The paper concludes that the future of designing these materials lies not in simply collecting more data, but in collecting the right kind of data that respects the complex physics of the system. By releasing their dataset, their experimental benchmark, and their trained models to the public, the researchers have provided a new standard for the field. This allows other scientists to design and screen new metal-organic frameworks with a level of confidence that was previously impossible, accelerating the discovery of materials that could help solve some of the world's most pressing energy and environmental challenges.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →