← Latest papers
🔬 materials science

Comparative study of ensemble-based uncertainty quantification methods for neural network interatomic potentials

This study systematically evaluates ensemble-based uncertainty quantification methods for neural network interatomic potentials and reveals that predictive precision often fails to reliably indicate accuracy in out-of-distribution regimes, frequently plateauing or decreasing as errors grow, thereby highlighting fundamental limitations in using uncertainty estimates as proxies for accuracy in large-scale materials simulations.

Original authors: Yonatan Kurniawan (Department of Physics and Astronomy, Brigham Young University, Provo, Utah, USA), Mingjian Wen (Institute of Fundamental and Frontier Sciences, University of Electronic Science and
Published 2026-07-01
📖 6 min read🧠 Deep dive

Original authors: Yonatan Kurniawan (Department of Physics and Astronomy, Brigham Young University, Provo, Utah, USA), Mingjian Wen (Institute of Fundamental and Frontier Sciences, University of Electronic Science and Technology of China, Chengdu, China), Ellad B. Tadmor (Department of Aerospace Engineering and Mechanics, University of Minnesota, Minneapolis, Minnesota, USA), Mark K. Transtrum (Cross Stream Consulting, Springville, UT, USA)

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: Teaching Computers to Guess How Atoms Behave

Imagine you are trying to teach a computer to predict how a material (like diamond or graphite) will behave. To do this, the computer needs a "rulebook" called an Interatomic Potential.

Traditionally, scientists used two types of rulebooks:

  1. The "Super-Accurate" Rulebook (Quantum Physics): It's incredibly precise but takes a supercomputer years to calculate just a few atoms. It's too slow for big projects.
  2. The "Fast & Simple" Rulebook (Classical Physics): It's very fast but often makes mistakes because it's too simple.

Machine Learning (ML) is the new middle ground. It's like a student who reads the "Super-Accurate" rulebook thousands of times and then learns to write its own "Fast & Simple" version that is almost as good but runs in seconds. These are called Machine Learning Interatomic Potentials (MLIPs).

The Problem: The "Confident Fool"

The main problem this paper tackles is trust.

When a student (the AI) learns from a textbook, they are great at answering questions that look exactly like the ones in the book. This is called being "In-Distribution" (ID). If you ask, "What is the energy of a diamond atom at room temperature?" and the AI saw that in its training, it will likely get it right.

But what happens if you ask a question the AI has never seen? Maybe you ask about a diamond under extreme pressure, or a strange shape it's never encountered? This is called being "Out-of-Distribution" (OOD).

The paper asks: Can the AI tell us when it is guessing?

In the world of AI, we use Uncertainty as a "confidence meter."

  • High Uncertainty: "I'm not sure about this answer. I might be wrong." (This is good! It warns us to be careful.)
  • Low Uncertainty: "I am 100% sure!" (This is good if the answer is right, but dangerous if the answer is wrong.)

The goal of the paper is to see if the AI's "confidence meter" actually works when the AI is forced to guess on new, weird data.

The Experiment: The "Study Group" Analogy

To test this, the researchers didn't just train one AI. They trained 100 different versions of the same AI. Think of this as a study group of 100 students who all studied the same textbook but learned slightly different things because they took notes in different ways.

They used four different ways to create this "study group" (Ensembles):

  1. Bootstrap: Giving each student a slightly different set of practice problems (resampling the data).
  2. Dropout: During the test, randomly telling some students to "shut up" (ignore parts of their knowledge) so the group has to rely on different ideas.
  3. Random Initialization: Giving each student a different starting brain (random weights) before they even began studying.
  4. Snapshots: Taking a photo of one student's brain at different times during their study session and treating those photos as different students.

The researchers then asked all 100 students to predict the behavior of carbon atoms in three forms: Diamond, Graphite, and Graphene.

The Results: The "Confident Fool" Strikes Again

Here is what they found, broken down simply:

1. When the AI is in its "Comfort Zone" (In-Distribution)

When the questions were similar to what the AI studied, the "confidence meter" worked okay. If the AI was unsure, it was usually because the answer was actually hard. If it was confident, it was usually right. The different study groups (ensembles) performed similarly here.

2. When the AI is pushed to the Edge (Out-of-Distribution)

This is where things got weird and dangerous. The researchers asked the AI to predict what happens when you stretch or squeeze the materials to extremes (like squeezing a diamond until it's tiny).

  • The AI got the answer wrong. (The predictions didn't match reality).
  • The AI said, "I am 100% sure I'm right!" (The uncertainty meter stayed low).

This is the "Confident Fool" problem. The AI is confidently making up facts.

3. The Counter-Intuitive Twist

Usually, you would expect that as the AI gets further away from what it knows, its "confidence meter" should go up (it should say, "Whoa, this is weird, I'm not sure!").

But the paper found the opposite.
As the AI was pushed further and further into the unknown (extreme stretching or squeezing), its confidence didn't go up. In fact, it sometimes went down or stayed flat.

  • Analogy: Imagine a driver who has only ever driven on a straight highway. If you suddenly ask them to drive on a steep, icy mountain pass, a smart driver would say, "I don't know how to do this!"
  • The AI Driver: Instead, the AI driver says, "I'm an expert!" and keeps driving at 100 mph, even though it's about to crash. The "uncertainty" actually decreased as the situation got more dangerous.

Why Does This Happen?

The researchers tried to figure out why the AI gets so confident when it's wrong. They had a few theories:

  • The "Saturated" Brain: The AI uses a mathematical function (like a switch that can only be fully "on" or fully "off"). When the input gets too extreme, the switch just stays stuck at "on," making the AI think it knows the answer perfectly, even though it's just guessing.
  • The "Groupthink" Effect: Even though they had 100 different AI models, when faced with a totally new situation, they all started making the same wrong guess. Because they all agreed with each other, the system thought, "Well, if 100 of us agree, we must be right!"

The Bottom Line

The paper concludes that we cannot blindly trust the AI's "confidence meter" when it is making predictions about things it hasn't seen before.

  • If the AI says, "I'm not sure," it's probably a good warning.
  • But if the AI says, "I'm 100% sure," it might be lying. It could be confidently predicting something completely wrong.

The authors warn that scientists need to be very careful when using these AI tools for big, new experiments. Just because the computer says it's accurate doesn't mean it is, especially when the computer is being asked to do something it wasn't trained for.

Key Takeaway: In the world of AI materials science, precision (confidence) does not equal accuracy (truth). You can be very precise and completely wrong at the same time.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →