The Linear Geometry of Interpretable Tokens: Jailbreaking Attacks and Defenses for Unlearned Diffusion Models
This paper reveals that erased concepts in unlearned diffusion models persist as coherent linear subspaces within token embeddings, leading to the development of SubAttack, a powerful jailbreaking method that exploits this structure, and SubDefense, a lightweight mechanism that projects out the subspace to robustly suppress residual harmful concepts while preserving generation quality.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where computers can paint pictures just by listening to your words. You say, "a cat wearing a hat," and poof, a perfect image appears. These are called Diffusion Models, and they are like digital artists that learned to create by studying millions of photos. But sometimes, these artists get a little too good at remembering things they shouldn't. If you ask them to draw something dangerous or copyrighted, they might do it anyway. To fix this, scientists invented a process called "Unlearning." Think of it like a teacher trying to make a student forget a specific fact, like the capital of a country, without making them forget how to do math or read. The goal is to "erase" the bad idea from the computer's brain while keeping its ability to draw everything else. But here's the tricky part: even after the teacher says, "Forget that fact," the student might still have a secret way of remembering it, hidden in a way we can't easily see. This paper dives into that mystery, asking: If we think we've erased a bad idea, where does it actually hide, and how can we finally get rid of it?
The researchers in this paper discovered that when a diffusion model tries to "unlearn" a harmful concept (like nudity or a specific artist's style), it doesn't actually delete it. Instead, the concept hides in a secret, linear subspace within the computer's memory. To use a metaphor, imagine the computer's brain is a giant library of words. When we try to remove the book about "Nudity," the computer doesn't burn the book. Instead, it shreds the pages and hides the scraps inside a specific, invisible drawer in the library. The paper shows that these scraps are still organized in a neat, predictable line. The authors call this hidden drawer a "residual subspace."
To prove this, the team created a new tool called SubAttack. Instead of guessing random words to trick the computer (which is how previous hackers tried), SubAttack acts like a master librarian who knows exactly where that invisible drawer is. It learns a special set of "magic keys" (which are just combinations of normal words) that unlock the drawer. When they used these keys, the "unlearned" computer immediately started drawing the harmful images again, proving the concept was never truly gone. What's fascinating is that these keys are made of words we can understand. For example, to make the computer draw a "nude" person again, the keys might be a mix of words like "slave," "hips," and "skin." This shows that the computer is still connecting the bad idea to these related, hidden clues.
The paper also found that this secret drawer isn't just in one computer; it's in many different "unlearned" models. If you find the key for one model, it often works on others, too. This suggests that the way these computers "forget" is flawed in a very similar way across the board. They aren't deleting the memory; they are just hiding it in a way that follows a straight, predictable line.
But the story doesn't end with the hack. Because the researchers understood exactly how the concept was hiding, they built a shield called SubDefense. If SubAttack is the key that opens the drawer, SubDefense is a heavy metal plate that welds that drawer shut. It works by taking the computer's entire library of words and mathematically "projecting" them away from that secret drawer. It's like telling the librarian, "No matter what word you use, make sure it never points toward that specific hidden drawer."
The results were impressive. When they used SubDefense, the computer became much harder to trick. Even when hackers tried to use the old methods or the new SubAttack keys, the computer refused to draw the harmful images. At the same time, the computer didn't lose its ability to draw safe, beautiful pictures. In fact, the images it drew after the defense were just as high-quality as before. The paper suggests that while we can't yet erase the concept perfectly, we can at least block the specific path it uses to sneak back in. This gives us a new, clear way to understand why "unlearning" is so hard and offers a practical, plug-and-play fix to make these digital artists safer for everyone.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.