← Latest papers
🔬 materials science

Bounding Retraining Equivalence and the Deletion Floor in Materials Machine Unlearning

This paper addresses the ambiguity of post-deletion accuracy in materials machine learning by defining a "deletion floor" based on retraining equivalence, establishing theoretical bounds that link low deletion floors to retained data redundancy, and demonstrating empirically that request-level unlearning evaluations must jointly report reference loss, prediction change, and retained utility to distinguish between effective unlearning and mere data suppression.

Original authors: Can Polat, Mustafa Kurban, Erchin Serpedin, Hasan Kurban

Published 2026-09-29
📖 5 min read🧠 Deep dive

Original authors: Can Polat, Mustafa Kurban, Erchin Serpedin, Hasan Kurban

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of materials science, researchers use powerful computer models to predict the properties of new substances before they are ever built in a lab. These models learn from vast databases of known chemical structures, mapping the arrangement of atoms to outcomes like how much energy a material stores or how stable it is. Because the laws of physics are consistent, many different chemical formulas can produce nearly identical crystal structures. This creates a natural redundancy in the data: if one specific record of a crystal is removed from the training set, the model can often still predict its properties accurately because it has learned the same physical patterns from other, very similar records that remain. This poses a unique challenge for a field called machine unlearning, which aims to make artificial intelligence systems "forget" specific data points, often to comply with privacy laws. If a system can still predict a removed item perfectly because of its neighbors, has it truly forgotten, or is it just using a different path to the same answer?

A team of researchers set out to solve this ambiguity by defining a new standard for what it means to successfully unlearn a piece of data in materials science. They introduced a concept called the "deletion floor," which represents the lowest possible error a model can have on a removed item if it were simply retrained from scratch without that item. Their work reveals that in many cases, the best possible retraining still predicts the removed item with high accuracy, simply because the remaining data is so similar. They found that the success of an unlearning attempt cannot be judged by accuracy alone; instead, it must be measured by comparing the model's new predictions against this specific baseline of what retraining would naturally produce.

The researchers tested these ideas using a massive collection of over 44,000 crystal structures from a public database known as the Materials Project. They focused on predicting formation energy, a measure of how much energy is required to build a crystal from its elements. To understand how redundancy affects forgetting, they first created a controlled experiment with synthetic data. They built groups of nearly identical items, then removed one member from each group and retrained the model. When they removed an item that had no close relatives, the model's error on that item remained higher than when they removed an item that had similar relatives remaining. In fact, the median error fell by nearly eight times when just one similar record was left behind. This demonstrated that the model was not truly "forgetting" the physical behavior of the removed crystal; it was simply relying on the surviving relatives to fill in the gap.

To see if this held up in real-world scenarios, the team applied their methods to the thousands of real crystal structures. They compared two different ways of training the models: one with a simpler setup and another with a more complex setup that used many more features to describe the crystals. They discovered a surprising trade-off. The more complex model was generally more accurate overall, but it was also much more sensitive to the removal of a single record. When a record was deleted, the complex model's prediction for that specific item shifted much more than the simpler model's prediction did. Yet, paradoxically, the complex model still ended up with a lower error rate on the deleted item after retraining. This means the model was highly sensitive to the change in its training data, yet it still managed to predict the removed item's value almost perfectly because the surrounding data was so informative.

The study also examined various methods used to approximate the unlearning process without actually retraining the model from scratch, a task that is computationally expensive. They found that some methods successfully increased the error on the specific deleted item, but often at the cost of making the model worse at predicting other, unrelated items. Others kept the overall model performance high but failed to significantly change the prediction for the deleted item, essentially leaving the "memory" of that record intact. The researchers concluded that simply checking if a model is less accurate on a deleted item is not enough. A successful unlearning process must be evaluated by looking at three things together: how much the prediction for the deleted item changed, how that change compares to what a full retraining would have produced, and whether the model's ability to predict other materials was preserved.

By establishing this "deletion floor," the researchers provided a clear reference point for the scientific community. They showed that in fields like materials science, where physical laws create strong connections between different data points, a model can appear to forget a record while still predicting it accurately. This is not a failure of the unlearning process, but a reflection of the underlying physics. The paper argues that future evaluations should stop treating post-deletion accuracy as a simple pass or fail metric. Instead, they should report the specific baseline error left by retraining, the actual shift in the model's prediction, and the overall health of the model. This approach ensures that when a system is asked to forget, we know exactly what it has let go of and what it has kept, grounded in the reality of the data rather than a vague notion of privacy.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →