When Are AI Explanations Recoverable? Identifiability, Stability, and Transfer from an Inverse-Problem Perspective
This paper establishes a mathematical framework for determining the recoverability of AI explanations by formulating them as inverse problems, deriving a complete trichotomy of identifiability and stability conditions for linear and nonlinear systems, and providing explicit bounds that distinguish intrinsic underdetermination from algorithmic variability.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of artificial intelligence, a model is often judged by how well it predicts the future. If a system can accurately forecast a patient's health outcome or a stock's price, we tend to trust its judgment. But trust in a prediction is different from trust in an explanation. When a model makes a decision, we often want to know why. Did it rely on a specific symptom? Did it weigh one factor more heavily than another? These explanations are meant to reveal the inner logic of the machine, acting as a window into its reasoning. However, a troubling reality exists: two different models can make the exact same predictions for every single case, yet offer completely different reasons for doing so. One might blame a specific feature while the other ignores it entirely, even though both arrive at the same correct answer. This creates a fundamental puzzle: if the observable behavior is identical, can we ever be sure which explanation is the true one?
This question sits at the heart of a new study that treats the search for AI explanations not as a matter of better software, but as a problem of physics and geometry. The researcher, Chibuike Chiedozie Ibebuchi at Morgan State University, approached the issue by framing it as an inverse problem. In science, a forward problem asks what happens when you push a button; an inverse problem asks what must have happened inside the machine to produce the result you see. The study asks whether the "why" of a prediction can be uniquely and stably recovered from the "what" of the prediction. The answer, the author finds, is not a simple yes or no. Instead, the recoverability of an explanation depends entirely on the mathematical shape of the data the model was trained on. The study proves that for some types of questions, the answer is hidden in plain sight, while for others, the answer is mathematically impossible to find, no matter how much computing power you apply.
The researcher developed a precise way to measure the limits of explanation. They imagined a scenario where two AI systems are so similar that their predictions differ by only a tiny, almost invisible amount. They then asked: how much can the explanations for these two systems differ? If the explanations can swing wildly even when the predictions are nearly identical, the explanation is unstable. If the explanations can be completely different even when the predictions are exactly the same, the explanation is unidentifiable. The study establishes a clear boundary between these states. It shows that an explanation is only recoverable if it aligns with the directions in the data that the model can actually "see." If the explanation depends on a direction the model cannot see, the explanation is lost forever. If the explanation depends on a direction the model can see but only faintly, the explanation exists in theory but is so sensitive to noise that it is useless in practice.
In the specific case of linear models, which are common in many statistical applications, the author derived an exact formula that acts as a ruler for this uncertainty. This formula links the stability of an explanation directly to the geometry of the data's relationships. When the data features are highly correlated, a situation known as collinearity, the ruler shows that explanations become increasingly unstable. The study demonstrates that as the correlation between data points grows, the uncertainty in the explanation grows with it, eventually becoming infinite if the correlation becomes perfect. The researcher tested this theory with computer simulations. They created scenarios where the data relationships were tweaked from perfectly clear to nearly impossible to distinguish. In every case, the computer's behavior matched the mathematical prediction exactly. When the data was well-behaved, the explanations were stable. When the data was nearly identical in different ways, the explanations diverged wildly, confirming that the instability was not a bug in the software but a fundamental property of the information available.
The study also explored what happens when the AI uses more complex, non-linear rules to make decisions. One might hope that a more complex model could untangle the confusion caused by messy data. However, the researcher found that complexity does not rescue a lost explanation. If a piece of information is missing from the data in the first place, no amount of non-linear twisting can bring it back. The study proves that if an explanation depends on a hidden variable that the model cannot observe, a non-linear model will still fail to identify it. The explanation remains just as invisible as it was in the simpler linear case. This finding is crucial because it dispels the idea that more complex algorithms can solve problems that are fundamentally underdetermined by the data.
Furthermore, the research addresses the issue of transfer, asking whether an explanation that works in one situation will hold up if the data changes slightly. The author found that stability is not guaranteed by small changes in the data. If the underlying data distribution shifts even a tiny bit, an explanation that was once stable can suddenly become impossible to recover. This happens when the data shift pushes the problem toward a boundary where information is lost. The study provides a specific condition, a "spectral margin," that must be maintained to ensure that explanations remain reliable when moving from one dataset to another. Without this margin, even a minuscule change in the data can destroy the ability to explain the model's behavior.
The work concludes by reframing how we should approach explainable artificial intelligence. Instead of immediately searching for the best algorithm to generate an explanation, the study suggests we must first ask if an explanation is even possible to find. Before trusting a feature attribution or a sensitivity score, we must determine if the information required to generate it is actually present in the observable predictions. The study separates the variability caused by the choice of algorithm from the intrinsic uncertainty caused by the data itself. It offers a mathematical basis for knowing when an explanatory claim is supported by the evidence and when it is fundamentally underdetermined. By defining the limits of what can be known, the research provides a necessary foundation for building trust in artificial intelligence, ensuring that we do not mistake a stable algorithm for a stable truth.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.