← Latest papers
🔬 materials science

Assessing the Transferability of General-Purpose MachineLearning Interatomic Potentials for Heterogeneous Catalysis with HetCat26

This paper introduces the HetCat26 benchmark to evaluate fifteen general-purpose foundation machine learning interatomic potentials for heterogeneous catalysis, revealing that while current models often fail to predict adsorption and DFT site preferences despite good performance on surface energetics and reaction barriers, their transferability cannot be inferred from general materials benchmarks and is significantly hindered by inconsistencies in training data.

Original authors: Alexandre Peuch, Giaan Kler-Young, Kaifeng Niu, Jinwoo Hwang, Manos Mavrikakis, Angelos Michaelides, Fabian Berger

Published 2026-09-28
📖 4 min read☕ Coffee break read

Original authors: Alexandre Peuch, Giaan Kler-Young, Kaifeng Niu, Jinwoo Hwang, Manos Mavrikakis, Angelos Michaelides, Fabian Berger

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Catalysts are the unsung heroes of modern industry, the materials that make chemical reactions happen faster and more efficiently without being used up themselves. They are essential for everything from making fertilizers to producing clean fuels, yet designing a better one is a slow, expensive process. For decades, scientists have relied on a powerful computer method called density functional theory to understand how atoms behave on a catalyst's surface. This method is incredibly accurate, treating the electrons in atoms with great precision, but it is also so computationally heavy that it can only simulate tiny systems for very short moments. To study real-world catalysts, which involve complex surfaces and reactions over longer times, researchers have turned to machine learning. These artificial intelligence models learn from the expensive computer simulations and then predict how atoms will behave much faster, acting as a bridge between high accuracy and practical speed. The hope has been that these new "foundation" models, trained on vast libraries of general chemical data, could be used immediately for any catalytic problem without needing to be retrained from scratch.

A team of researchers set out to test whether these general-purpose models are truly ready for the specific, demanding world of heterogeneous catalysis, where reactions happen at the interface between a solid surface and a gas or liquid. They created a new set of tests called HetCat26, designed to mimic the real challenges a catalyst faces: how atoms stick to a surface, how they move around, and how they break apart to form new chemicals. They ran fifteen different pre-trained machine learning models through this gauntlet to see how well they performed compared to the gold-standard computer simulations. The results revealed a surprising truth: doing well on general chemistry tests does not guarantee success in catalysis. While some models handled the energy of the surface itself quite well, many struggled significantly with how molecules attach to that surface, a critical step for any reaction to occur.

The study found that the models were surprisingly good at predicting the energy required to break bonds during a reaction, a task that involves atoms in unstable, high-energy states. For example, when simulating the conversion of carbon dioxide into methanol or the water-gas shift reaction, the best models could predict the energy barriers with an error of less than five percent, a level of accuracy that rivals custom-built tools. They also did a decent job describing the energy of the metal surfaces themselves. However, the models faltered when it came to adsorption, the process where a molecule lands and sticks to the catalyst. In many cases, the models could not correctly predict which specific spot on the surface a molecule would choose to bind to, or they miscalculated the strength of that bond by a large margin. This is a major hurdle because if a molecule binds too tightly, it poisons the surface; if it binds too weakly, it bounces off before reacting.

Perhaps the most significant finding was that a model's performance on broad, general materials benchmarks was a poor predictor of its performance in catalysis. A model that ranked as a top performer for general material science could be one of the worst performers for catalytic reactions. This suggests that the unique, complex environments found at catalyst surfaces are not well represented in the general training data. The researchers also discovered that mixing different types of computer calculations in the training data caused confusion. Specifically, combining data calculated with one standard method and data calculated with a modified method for certain metals led to large errors. When they removed these inconsistent data points, the models performed much better.

Two models, eSEN-30M-OAM and MACE-MH-1-OMAT, stood out as particularly promising, achieving high accuracy across almost all the tests without needing any special tuning. These models managed to capture the subtle energy differences between different atomic arrangements and reaction steps with remarkable precision. However, even these top performers showed weaknesses when dealing with metal clusters on oxide surfaces or highly irregular atomic sites, which are common in real-world catalysts but rare in the training data. The researchers concluded that while foundation models are powerful tools, they are not yet a universal solution for catalysis. To make them truly reliable, future training sets must be more consistent in their underlying calculations and must include more examples of the specific, tricky environments found at the heart of catalytic reactions. This new benchmark, HetCat26, provides a clear roadmap for the next generation of models, ensuring that the artificial intelligence guiding the design of future catalysts is built on a foundation of accurate and relevant data.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →