Shape of Memory: a Geometric Analysis of Machine Unlearning in Second-Order Optimizers
The paper argues that current machine unlearning definitions are insufficient for second-order optimizers because, while they appear to achieve performance goals, they retain detectable residual information within their optimizer states that can only be fully erased through controlled geometric perturbations.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are training a professional chef to cook a signature dish.
Most people think "learning" is just about the chef memorizing the final recipe (the parameters). If you want the chef to "unlearn" a specific ingredient—say, they accidentally learned to use salt when they should have used sugar—most current methods just tell the chef, "Hey, change the final recipe; don't use salt."
But this paper argues that learning is much deeper than just the recipe. It’s about the chef’s intuition (the second-order optimizer state).
The Core Problem: The "Muscle Memory" of Math
When a chef learns, they don't just learn the ingredients; they learn the rhythm of the kitchen. They learn how heavy the knife is, how much pressure to apply to the garlic, and how the heat of the stove reacts to different oils. This "rhythm" is the second-order state. It’s a complex, invisible map of how every action affects the next.
Current "unlearning" techniques only fix the recipe. They change the final output so the dish tastes right. But they forget to fix the chef's intuition.
The paper’s big discovery is this: Even if the chef produces a perfect, salt-free dish (the model performs well), their rhythm is still "salty." Their muscle memory is still shaped by that mistake. If you watch them work, you can see the "ghost" of the deleted information in the way they move their hands.
The "Ghost in the Machine" (Volatility)
The researchers looked at two types of learners:
- The Simple Learner (First-Order): This is like a chef who only follows a checklist. If you tell them to stop using salt, they stop. They don't have much "intuition" to mess up. They recover quickly.
- The Intuitive Learner (Second-Order): This is the high-level chef. They use "curvature" (the shape of the kitchen) to work faster. When you tell them to "unlearn" an ingredient, they can fix the recipe easily, but their internal geometry—their sense of rhythm—gets jittery and weird.
The researchers found that even when the "Intuitive Learner" starts making correct predictions again, their internal state shows "volatility." It’s like a person who has been told to forget an ex-partner; they might act perfectly normal in public, but their heart rate spikes or their hands shake whenever a certain song plays. The information is "gone" from the output, but it’s still vibrating in the system.
The Solution: "Geometric Erasure"
The paper suggests that if we want to truly unlearn something, we can't just change the recipe. We have to perform a "geometric perturbation."
In our kitchen analogy, this isn't just telling the chef "don't use salt." It’s like slightly rearranging the kitchen layout or changing the weight of the knives. By shaking up the chef's physical environment (the optimizer's state), you force them to break their old, "salty" muscle memory and build a brand-new rhythm from scratch.
Summary in a Nutshell
- The Old Way: "Change the answer so the mistake isn't visible."
- The Paper's Way: "The mistake is still hiding in the way the machine thinks. To truly forget, you have to shake up the machine's internal sense of direction, not just its final answer."
The takeaway: If we want true privacy (making sure data is truly deleted), we can't just look at what the AI says; we have to look at how the AI thinks.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.