← Latest papers
🔢 mathematics

Optimal information deletion and Bayes' theorem

This paper revisits Arnold Zellner's seminal work on Bayes' theorem by demonstrating that the optimal rule for deleting information—updating a posterior to an antedata distribution without creating or destroying information—coincides with the leave-data-out posterior derived from Bayes' theorem.

Original authors: Hans Montcho, Håvard Rue

Published 2026-05-20
📖 4 min read🧠 Deep dive

Original authors: Hans Montcho, Håvard Rue

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart detective who has solved a complex case using a massive pile of evidence. This detective has written a final report (the posterior distribution) that perfectly combines everything they know about the suspect (the prior) with all the clues they found (the data).

Now, imagine a situation where you need to "un-solve" part of the case. Perhaps a specific piece of evidence turns out to be a mistake, or a user demands that their personal data be completely erased from the detective's memory. You need the detective to produce a new report that looks exactly like the one they would have written if they had never seen that specific piece of evidence in the first place.

This is the problem of Bayesian Unlearning (or data deletion).

The Old Way vs. The New Insight

Usually, if you want to remove a piece of evidence, the easiest way is to throw away the whole report and start over from scratch, re-reading every single clue without the bad one. But if the detective has analyzed millions of clues, starting over is slow and expensive.

In 1988, a mathematician named Arnold Zellner proved that Bayes' Theorem is the most efficient way to add information. He showed that when you take a guess and add new clues, the standard mathematical formula for updating your guess is the only way to do it without accidentally inventing fake facts or losing real ones.

This new paper, written by Hans Montcho and Håvard Rue, asks the "backward" question: Is Bayes' Theorem also the most efficient way to remove information?

The "Perfect Eraser" Analogy

The authors treat information like a physical substance. They propose a rule for "Information Deletion."

  • The Goal: Take the detective's final report and surgically remove the influence of a specific clue (ygy_g) to get a report based only on the remaining clues (ygy_{-g}).
  • The Constraint: The rule must be "optimal." This means it cannot destroy real information (making the detective forget things they should remember) and it cannot create fake information (hallucinating new facts that weren't there).

The authors use a mathematical tool called "variational calculus" (think of it as a way to find the smoothest, most perfect path through a maze) to find the best rule for this deletion.

The Big Discovery

After doing the math, the authors prove a surprising and elegant result:

The "Optimal Information Deletion Rule" is exactly the same as the standard Bayes' Theorem.

In other words, the mathematical operation you use to add a clue is the exact same operation (just reversed) that you use to remove a clue.

  • Adding Info: You take your current belief and multiply it by the new clue.
  • Removing Info: You take your current belief and divide it by the clue you want to forget.

The paper proves that if you try to invent a shortcut to delete data, you will inevitably either lose too much information or accidentally create new, false information. The only way to delete data perfectly, without distortion, is to use the standard Bayes' formula in reverse.

Why This Matters (According to the Paper)

The authors highlight two main takeaways:

  1. Symmetry: Just as the French chemist Lavoisier said "Matter is neither created nor destroyed," the authors suggest that in Bayesian learning, information is neither created nor lost; it is just transformed. Bayes' Theorem is the perfect tool for both learning (adding) and unlearning (removing).
  2. Better Approximations: Since we know the "perfect" way to delete data is to divide by the likelihood of that data, we can now build better "approximate" methods. If we can't do the perfect math (because it's too hard), we can use this new understanding to build shortcuts that are guaranteed to be the best possible approximations, rather than just random guesses.

Summary

Think of the detective's report as a cake.

  • Learning is adding a new layer of frosting (data) to the cake.
  • Unlearning is trying to scrape off that specific layer of frosting to see the cake underneath.

This paper proves that the only way to scrape off the frosting without ruining the cake underneath is to use the exact same tool you used to put the frosting on, just in reverse. Any other tool will either leave crumbs behind (lost information) or scrape off part of the cake itself (destroyed information).

The paper concludes that Bayes' Theorem is the "Goldilocks" rule: it is perfectly balanced for both adding and removing information, ensuring that the truth remains intact throughout the process.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →