← Latest papers
💬 NLP

Beyond Forgetting: Machine Unlearning Elicits Controllable Side Behaviors and Capabilities

This paper demonstrates that Representation Misdirection, a machine unlearning method for large language models, not only achieves forgetting but also elicits controllable side behaviors and enhances capabilities by redirecting latent representations toward a target vector aligned with high-level concepts.

Original authors: Tien Dang, The-Hai Nguyen, Dinh Mai Phuong, Nguyen Minh Phuong, Anh Bui, Hoang Thanh-Tung, Le-Minh Nguyen, Naoya Inoue

Published 2026-05-18
📖 4 min read☕ Coffee break read

Original authors: Tien Dang, The-Hai Nguyen, Dinh Mai Phuong, Nguyen Minh Phuong, Anh Bui, Hoang Thanh-Tung, Le-Minh Nguyen, Naoya Inoue

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a giant, super-smart library (a Large Language Model) that knows everything. Sometimes, you need to remove specific books from this library because they contain dangerous or private information. This process is called "Machine Unlearning."

For a long time, the way researchers tried to remove these books was a bit like throwing a bucket of random noise over the shelves. They would tell the library, "Forget this specific topic!" by scrambling the memory of those books with static. The problem? This often made the whole library go crazy. It might start speaking gibberish, losing its ability to tell the truth, or refusing to answer harmless questions.

This paper introduces a smarter, more surgical way to do this, and it discovers a surprising side effect: You can actually program the library to behave in specific new ways while you are deleting the old stuff.

Here is how the paper explains it, using simple analogies:

1. The "Linear Compass" Idea

The authors rely on a theory called the Linear Representation Hypothesis. Imagine that inside the library's brain, every big idea (like "Truth," "Sadness," or "Refusal") has a specific direction, like a compass needle pointing North.

  • If you want the library to be more truthful, you just need to push its thoughts slightly toward the "Truth" direction.
  • If you want it to be more negative, you push it toward the "Sadness" direction.

2. The Two New Tools: "Addition" and "Ablation"

The paper proposes two ways to use these compass directions to delete information:

  • Representational Addition (RAd): Imagine you are trying to delete a dangerous book about "How to build a bomb." Instead of just erasing the page, you take the library's memory of that book and push it hard in the direction of "Truth."
    • The Result: The library forgets the bomb instructions (because it's now focused on being truthful), but as a bonus, the library becomes better at telling the truth in general. It didn't just forget; it learned a new superpower aligned with the direction you pushed it.
  • Representational Ablation (RAb): This is the opposite. Imagine you want to delete a book about "Refusing to help." You take the library's memory and project it sideways, effectively cutting out the part of the memory that says "No."
    • The Result: The library forgets the refusal instructions, but as a side effect, it becomes less likely to refuse even when it should. It loses that specific "No" button.

3. The Surprising Discovery: "Controllable Side Effects"

The biggest finding of the paper is that this isn't just about deleting. It's about steering.

By choosing which direction you push the memory toward, you can make the "unlearned" model do cool new things:

  • Truthfulness: If you push the memory toward the "Truth" direction, the model becomes much better at answering tricky questions correctly.
  • Sentiment: If you push it toward "Positive," the model starts sounding happier. If you push it toward "Negative," it sounds sadder.
  • Language: If you push it toward "French," the model starts answering your English questions in French!
  • Reasoning: If you push it toward "Step-by-Step thinking," the model gets better at solving math problems without you having to teach it new math.

4. Why Randomness Was the Problem

Previous methods used a random direction (like pointing the compass at a random spot on the map).

  • The Paper's Claim: If you push the memory in a random direction, the model just gets confused or loses its general smarts.
  • The New Way: If you push it toward a specific concept (like "Truth" or "French"), the model forgets the bad stuff and gains that specific skill.

5. Is it a Risk or a Feature?

The authors warn that this is a double-edged sword:

  • The Risk: If someone uses this to make a model forget safety rules, they could accidentally (or intentionally) make the model more likely to say "No" to harmless things or become overly aggressive.
  • The Feature: If used correctly, you can create a model that has forgotten dangerous secrets but is now better at being honest, speaking a specific language, or solving logic puzzles.

Summary

Think of Machine Unlearning not as a "delete button," but as a tuning knob.

  • Old way: Turn the knob randomly to delete, and hope the radio doesn't break.
  • New way (this paper): Turn the knob toward a specific station (like "Truth" or "French"). You delete the unwanted noise, but you also tune the radio to play that specific song perfectly.

The paper proves that by understanding the "compass directions" inside AI, we can delete bad knowledge while simultaneously upgrading the AI's other skills.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →