← Latest papers
⚡ electrical engineering

Physics-guided explainable machine learning and uncertainty quantification for predicting unconfined compressive strength of nano shale–quicklime stabilized kaolin clay

This study proposes a physics-guided, explainable machine learning framework using an optimized CatBoost model to accurately predict the unconfined compressive strength of nano shale–quicklime stabilized kaolin clay while quantifying uncertainties and identifying key physical drivers like plasticity and curing interactions.

Original authors: Arash Aminaee, Abolfazl Soltani, Abolfazl Baghbani, Parisa Salehan, Hossam Abuel-Naga, Pijush Samui

Published 2026-09-12
📖 7 min read🧠 Deep dive

Original authors: Arash Aminaee, Abolfazl Soltani, Abolfazl Baghbani, Parisa Salehan, Hossam Abuel-Naga, Pijush Samui

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The ground beneath our feet is rarely uniform. In many places, the soil is soft, wet, and prone to shifting, making it a poor foundation for roads, buildings, or bridges. Engineers have long known how to fix this by mixing the weak earth with stabilizing agents like lime or cement. These additives trigger chemical reactions that bind soil particles together, turning a squishy mess into a solid, load-bearing material. However, figuring out exactly how much stabilizer to use and how long to wait for the soil to harden is a slow, expensive process. It requires mixing dozens of different batches in a lab, waiting weeks for them to cure, and then crushing them to measure their strength. For every new type of soil or every change in the recipe, this cycle must be repeated, consuming time and resources that could be better spent on building.

A team of researchers has developed a new way to speed up this process without sacrificing accuracy. They created a computer model that acts as a highly skilled predictor, learning from past experiments to guess the strength of future soil mixtures. But instead of just feeding the computer raw numbers, the researchers taught it to think like a geotechnical engineer. They built the model using "physics-guided" features, which means they gave the computer descriptors that reflect the actual physical and chemical processes happening in the soil, such as how the binder interacts with time and how the soil's natural plasticity changes. This approach allows the model to understand the "why" behind the strength, not just the "what." The result is a tool that can screen potential soil mixtures quickly and reliably, helping engineers decide which combinations are worth testing in the lab and which are likely to fail, all while providing a clear explanation of why a specific mixture is expected to perform well.

The researchers focused on a specific type of soil treatment involving kaolin clay, a common weak soil, stabilized with two materials: nano shale, a finely ground rock powder, and quicklime, a reactive form of calcium oxide. They started with a database of 180 individual laboratory measurements taken from 36 different mixture designs. Each design varied the amount of nano shale and lime used, as well as the number of days the soil was allowed to cure. The goal was to predict the unconfined compressive strength, a standard measure of how much pressure the soil can withstand before breaking. To do this, they tested several different types of machine learning algorithms, ranging from simple linear equations to complex artificial neural networks and advanced tree-based models. They did not simply split the data into random training and testing groups, a common practice that can sometimes hide flaws in a model. Instead, they used a rigorous method called nested group cross-validation. This ensured that the model was tested on entirely new mixture designs it had never seen before, preventing it from simply memorizing the results of similar samples and giving a false sense of accuracy.

After training and testing these models, the researchers found that a specific algorithm called CatBoost performed the best, but only after they refined the information it received. Initially, they gave the model seven engineered features designed to capture the physics of the stabilization process, such as the total amount of binder used and how that binder interacted with the curing time. Through a careful process of removing features one by one, they discovered that one of these inputs was actually redundant and was confusing the model. By removing it, they created a streamlined version with six key descriptors that included the total binder content, the interaction between the nano shale and the curing days, and a measure of the soil's relative plasticity. This refined model, which they named CatBoost-6PG, achieved a high level of accuracy, correctly predicting the strength of the soil mixtures with an error margin of about 151 kilopascals. This performance was significantly better than the simpler models and even outperformed the more complex ones that did not use these physics-based inputs.

The true value of this work, however, lies in its transparency. The researchers did not stop at getting a good prediction; they wanted to understand what the model was actually learning. Using a technique called SHAP analysis, they mapped out which factors mattered most. They found that the model correctly identified the soil's plasticity index—a measure of how much water the soil can hold before becoming sticky—as the most critical factor. High plasticity consistently led to lower strength predictions, which aligns perfectly with established soil mechanics. The model also learned that the passage of time was crucial, with longer curing periods leading to stronger soil, and that the interaction between the binder and the time allowed for curing was more important than the amount of binder alone. For instance, the model showed that adding more nano shale only helped if there was enough time for the chemical reactions to occur. These insights confirm that the model is not just finding statistical patterns but is capturing the real physical mechanisms of soil stabilization.

To ensure the predictions were trustworthy, the team also quantified the uncertainty, or the range of possible outcomes, for each prediction. They simulated thousands of scenarios where the input values, such as the exact amount of lime or the curing time, were slightly varied to mimic real-world imperfections. They found that the biggest source of uncertainty came from the variability in the input materials themselves, rather than from the model's internal calculations. When they accounted for this variability, the model provided a prediction interval that covered nearly 90 percent of the actual test results. This means that while the model gives a specific number for the expected strength, it also provides a safety margin, telling engineers how much that number might fluctuate in reality. This is a vital feature for engineering, where safety depends on understanding the limits of a prediction.

The study also explored how useful the model would be in the very early stages of a project, when engineers might not yet have detailed data on the soil's compaction properties or plasticity limits. They tested the model using only the basic mix design variables: the amounts of nano shale and lime, and the curing time. While the model could still make a rough estimate with this limited information, its accuracy dropped significantly. This finding is practical and honest: the model is not a replacement for laboratory testing, but a powerful tool to guide it. It can help engineers narrow down the list of promising mixtures to test, saving time and money, but it still relies on the fundamental soil properties to make its most accurate predictions.

Ultimately, this research offers a bridge between raw data and engineering intuition. By combining machine learning with a deep understanding of the physical processes at work, the researchers created a tool that is both accurate and explainable. It does not treat the soil as a black box but instead learns the rules of how binders, time, and soil properties interact to create strength. This approach allows for a more efficient design process, where engineers can confidently screen mixtures and focus their resources on the combinations most likely to succeed. The work demonstrates that when artificial intelligence is guided by the laws of physics, it can become a reliable partner in solving complex engineering challenges, offering clarity and confidence in the design of the structures that support our daily lives.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →