← Latest papers
💻 computer science

Erasure or Erosion? Evaluating Compositional Degradation in Unlearned Text-To-Image Diffusion Models

This paper systematically evaluates post-hoc unlearning in text-to-image diffusion models and reveals a critical trade-off where effective removal of undesirable concepts, such as nudity, frequently compromises the model's broader compositional generation capabilities like attribute binding and spatial reasoning.

Original authors: Arian Komaei Koma, Seyed Amir Kasaei, Ali Aghayari, AmirMahdi Sadeghzadeh, Mohammad Hossein Rohban

Published 2026-04-07
📖 4 min read☕ Coffee break read

Original authors: Arian Komaei Koma, Seyed Amir Kasaei, Ali Aghayari, AmirMahdi Sadeghzadeh, Mohammad Hossein Rohban

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a incredibly talented artist who can draw anything you describe. If you say, "a green banana and a brown dog," they can instantly paint a perfect picture. However, this artist learned from a massive library of books and magazines that included some things you don't want them to draw, like explicit nudity or copyrighted characters.

You can't afford to fire the artist and hire a new one (retraining the AI is too expensive), so you try to "unlearn" the bad habits. You tell the artist, "Never draw nudity again."

This paper asks a very important question: When you force the artist to forget the bad stuff, do they accidentally forget how to draw good stuff too?

The authors call this the battle between "Erasure" (successfully removing the bad) and "Erosion" (wearing down the artist's general skills).

The Experiment: The "Banana and Dog" Test

To test this, the researchers took several different "unlearning" techniques and applied them to a popular AI model (Stable Diffusion). They then gave the models a safe, innocent prompt: "Draw a green banana and a brown dog."

Here is what they found, using some simple analogies:

1. The "Over-Correction" Problem (The Erasure Trap)

Some methods were like a strict teacher who, when told "Don't draw nudity," decided the safest thing to do was to stop drawing anything interesting at all.

  • The Result: These models became "safe," but they also became broken. When asked for a green banana, they might draw a brown banana, or forget to draw the dog entirely, or put the dog inside the banana.
  • The Metaphor: Imagine a chef who is told, "Never use salt." To be safe, they decide to stop seasoning any food. Now, the steak tastes like cardboard, and the soup is flavorless. They successfully removed the "salt" (the bad concept), but they ruined the "flavor" (the model's ability to understand how things go together).
  • The Data: The most aggressive methods (like Scissorhands and EraseDiff) achieved perfect safety scores but destroyed the model's ability to understand spatial relationships (where things are) and attributes (what color things are).

2. The "Surgical" Approach (The Preservation Hope)

Other methods tried to be more like a surgeon, carefully removing only the specific bad concept without touching the rest of the brain.

  • The Result: These models (like SPM and ACE) were better at keeping their general skills. They could still draw a green banana and a brown dog correctly.
  • The Catch: They weren't always 100% perfect at removing the bad stuff. Sometimes, a tiny bit of the "forbidden" concept might slip through.
  • The Metaphor: This is like a chef who carefully removes the salt shaker but keeps the pepper and spices. The food might still have a tiny pinch of salt if they aren't careful, but the steak still tastes delicious.

The Big Trade-Off

The paper reveals a painful truth: You generally can't have both.

  • If you demand 100% safety (perfect erasure), the AI often loses its "common sense." It forgets how to bind colors to objects or count items. It's like a person who, after being told to forget a specific word, starts speaking in gibberish because they forgot how grammar works.
  • If you demand the AI keep its skills (perfect composition), it often fails to completely erase the bad concept.

Why This Matters

The authors argue that we are currently measuring success the wrong way. We are only checking, "Did the AI draw the nudity?" If the answer is "No," we say, "Great job!"

But the paper says, "Wait, look at the banana!" If the AI can't draw a green banana anymore because we forced it to forget nudity, then the AI is broken. It's not just "safe"; it's "semantically damaged."

The Takeaway

The paper concludes that we need to stop treating AI safety like a simple on/off switch. We need to teach AI to be safe without breaking its brain. A model that is technically safe but can no longer understand the world (e.g., it doesn't know that a banana is usually yellow or that a dog is usually on the ground) is not a useful tool.

In short: Don't just fire the bad habits; make sure the artist keeps their talent. If you erase too much, you don't get a better artist; you just get a blank canvas.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →