← Latest papers
🔬 materials science

A single design choice determines whether machine learning models of materials make physically impossible predictions

This paper demonstrates that incorporating parity labels into machine learning model features is a critical design choice that guarantees physically exact predictions for symmetry-forbidden properties in centrosymmetric crystals, whereas models lacking this feature produce physically impossible errors in nearly all cases without sacrificing accuracy.

Original authors: Can Polat, Mustafa Kurban, Erchin Serpedin, Hasan Kurban

Published 2026-08-20
📖 6 min read🧠 Deep dive

Original authors: Can Polat, Mustafa Kurban, Erchin Serpedin, Hasan Kurban

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the quest to discover new materials, scientists have long relied on a powerful shortcut: using the rules of symmetry to predict how a crystal will behave. Crystals are not random piles of atoms; they are highly ordered structures where atoms repeat in precise patterns. Because of this order, the laws of physics dictate that certain properties must behave in specific ways. For instance, if a crystal looks exactly the same when you turn it upside down through its center—a property known as centrosymmetry—then it cannot generate electricity when squeezed. This is not a matter of the material being weak or the measurement being imprecise; it is a hard rule of nature. If a crystal has this central symmetry, the electric response must be exactly zero.

For decades, researchers have used complex computer simulations based on quantum mechanics to calculate these properties. Today, however, a faster method is taking over: machine learning. These artificial intelligence models are trained to predict material properties by learning from vast databases of known crystals. They are becoming the primary tool for discovering new substances for batteries, solar cells, and electronics. The hope is that these models can learn the rules of symmetry just as a human physicist would, or perhaps even better, by finding patterns in the data that humans might miss. But this speed comes with a risk: if the model does not strictly obey the fundamental laws of symmetry, it might predict that a crystal can do something that nature forbids, such as generating electricity when it physically cannot.

A recent study by researchers at Texas A&M University and other institutions reveals a startling flaw in how many of these popular machine learning models are built. The team discovered that whether a model respects these absolute laws of physics is not determined by how much data it sees or how long it is trained. Instead, it is decided by a single, tiny design choice made before the training even begins. This choice is whether the model's internal "features"—the mathematical building blocks it uses to understand the shape of a crystal—carry a specific label called "parity." Parity is a way of tracking how a shape changes when it is reflected in a mirror or turned inside out. The researchers found that if a model lacks this label, it is mathematically incapable of predicting the exact zero required by symmetry, no matter how hard it tries. It will always produce a small, non-zero number, which represents a physically impossible prediction.

To prove this, the researchers set up a rigorous test using a property called the piezoelectric tensor, which measures how much electricity a crystal produces when squeezed. They focused on two thousand different crystals that are known to be centrosymmetric, meaning they have a center of symmetry. According to the laws of physics, the piezoelectric response for all of these crystals must be exactly zero. The team took several of the most advanced machine learning models used in the field and created two versions of each. One version was built with the parity labels included, while the other was built without them, identical in every other way. They then trained both versions on the same data and asked them to predict the piezoelectric response for the two thousand centrosymmetric crystals.

The results were dramatic and immediate. The models with the parity labels predicted a value of zero for every single crystal, down to the limits of the computer's precision. They correctly identified that these materials could not produce electricity. In stark contrast, the models without the parity labels failed spectacularly. They predicted a non-zero electrical response for 90 to 96 percent of the crystals. These predictions were not tiny errors; they were often as large as the real electrical responses found in materials that can generate electricity. The difference between the two groups was not a matter of a few decimal places; it was a gap of six orders of magnitude, meaning the "wrong" models were a million times more likely to predict a physically impossible event than the "right" models.

The researchers then tested whether training could fix the problem. They fed the "wrong" models extra data, specifically showing them examples of centrosymmetric crystals with the correct zero label, hoping the models would learn to respect the rule. It did not work. Even after seeing thousands of examples where the answer was zero, the models continued to predict non-zero values. They could get closer to zero, but they could never reach the exact zero required by the laws of physics. The study also showed that simply averaging the results of a model and its mirror image did not solve the issue, because the underlying structure of the model itself was blind to the symmetry rule.

This finding has profound implications for the field of materials science. It suggests that the reliability of a machine learning model is not just about its accuracy on standard tests, but about its structural ability to represent physical reality. The researchers audited eighteen different machine learning architectures currently in use and found that exactly half of them lack the necessary parity labels. This means that a significant portion of the tools scientists are using to discover new materials are fundamentally incapable of making the correct prediction for a large class of crystals. The problem is not that these models are bad at learning; it is that their design prevents them from learning this specific truth.

The study also highlighted a growing trend in the field where scientists take a pre-trained model, which has already learned the general structure of atoms, and attach a new "head" to it to predict a specific property. The researchers showed that if the original model was built without parity labels, the new property predictor will inherit that flaw. It does not matter how carefully the new part is designed; if the foundation is missing the symmetry label, the entire structure will fail to respect the laws of physics. This means that the decision to include or exclude this label is made upstream, often by the creators of the base model, and it silently dictates the validity of every prediction made by anyone who uses that model later.

Ultimately, the paper argues that for machine learning to be truly trustworthy in science, it must be built with the laws of physics baked into its structure, not just learned from data. The researchers propose a simple check: before trusting a model to predict a property that should be zero due to symmetry, one can test it in seconds by applying a mirror reflection to a random crystal structure. If the model predicts a non-zero value, it lacks the necessary parity label and cannot be trusted for that task. This is not a suggestion for improvement but a requirement for physical validity. The study concludes that the presence or absence of this single label determines whether a model is a tool for discovery or a source of impossible predictions, and it is a detail that must be reported and verified for every model used in the search for new materials.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →