← Latest papers
💻 computer science

Erased, But Not Forgotten: Erased Rectified Flow Transformers Still Remain Unsafe Under Concept Attack

This paper introduces ReFlux, a novel concept attack method that exploits attention localization in rectified flow transformers like Flux to demonstrate that current concept erasure techniques remain ineffective and unsafe against targeted reactivation.

Original authors: Nanxiang Jiang, Zhaoxin Fan, Enhan Kang, Daiheng Gao, Yun Zhou, Yanxia Chang, Zheng Zhu, Yeying Jin, Wenjun Wu

Published 2026-04-09
📖 5 min read🧠 Deep dive

Original authors: Nanxiang Jiang, Zhaoxin Fan, Enhan Kang, Daiheng Gao, Yun Zhou, Yanxia Chang, Zheng Zhu, Yeying Jin, Wenjun Wu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: The "Magic Eraser" That Wasn't Magic Enough

Imagine you have a super-smart artist (an AI) who can draw anything you describe. But this artist has learned some bad habits from the internet—they can draw inappropriate or harmful things if you ask for them.

To fix this, developers tried to use a "Magic Eraser" (Concept Erasure). They told the artist, "Forget how to draw nudity, violence, or specific celebrities." They thought they had wiped the artist's memory clean.

The Problem: This paper reveals that the "Magic Eraser" didn't actually wipe the memory. It just put a heavy blanket over the bad thoughts. The artist still remembers how to draw them; they just need someone to pull the blanket off.

The authors created a new tool called ReFlux (the "Blanket Puller") to prove that even after these safety measures, the AI is still unsafe.


The Characters in Our Story

  1. The Artist (Flux): This is the new, super-powerful AI model. Unlike older artists who drew one brushstroke at a time (like Stable Diffusion), Flux is like a symphony conductor. It looks at the whole picture and the words together, flowing smoothly from a blank canvas to a finished masterpiece.
  2. The Safety Team (The Erasers): These are the people who tried to "unlearn" bad concepts. They tried to surgically remove the specific "notes" in the AI's song that correspond to bad ideas.
  3. The Hackers (ReFlux): The authors of this paper. They aren't trying to make bad art; they are trying to test the safety. They want to see if the "unlearning" actually worked.

The Core Discovery: Why Old Tricks Failed

The researchers tried using old tricks (attacks designed for the older artists) to wake up the new artist (Flux), but they failed. Why?

  • Old Artist (Stable Diffusion): If you want to make the artist draw a "cat," you just whisper the word "cat" louder. The artist listens to individual words.
  • New Artist (Flux): This artist listens to the whole sentence and the flow of the music. It doesn't care about individual words as much; it cares about how the words connect to the image.

The researchers realized that the Safety Team's "Magic Eraser" worked by hiding the spotlight.

  • Imagine the AI has a spotlight that shines on the word "cat" in the prompt.
  • The Safety Team dimmed that spotlight so the artist couldn't see the word "cat."
  • The Flaw: The artist still knows where the word "cat" is; the spotlight is just turned down low.

The Solution: ReFlux (The Spotlight Booster)

The authors built ReFlux, a clever tool that doesn't just shout the word louder. Instead, it does three smart things:

  1. The Spotlight Booster (Attention Reactivation): It gently turns the dimmed spotlight back up, but carefully so it doesn't blind the artist or ruin the rest of the painting.
  2. The Flow Guide (Velocity Guidance): Since Flux is a "flow" artist, ReFlux nudges the artist's brushstrokes in the right direction, guiding the paint to form the erased concept without messing up the rest of the picture.
  3. The Memory Keeper (Consistency): It makes sure that while the "cat" reappears, the "dog" and the "background" stay exactly the same. It doesn't want to break the whole painting to fix one part.

The Result: With only a tiny amount of extra memory (like a small USB drive), ReFlux successfully made the "erased" concepts appear again, proving the Safety Team's work was incomplete.


The Analogy: The "Whispering Library"

Imagine the AI is a library with millions of books.

  • The Safety Team took all the books about "Dangerous Ideas" and locked them in a dark basement. They thought no one could find them.
  • Old Attackers tried to shout "DANGER!" at the library entrance, but the new library (Flux) has soundproof walls, so the shouting didn't work.
  • ReFlux realized the books weren't gone; they were just in the dark. ReFlux didn't shout; it simply turned on a flashlight in the basement and walked right up to the books.

The authors found that even though the books were locked away, the library staff (the AI) still knew exactly where they were and could pull them out if you showed them the flashlight.


Why Does This Matter?

You might ask, "Why are you trying to break the safety?"

The authors argue that you can't fix what you can't measure.

  • If we think the "Magic Eraser" works, we might let these AIs loose on the internet, thinking they are safe.
  • But if they are actually just "hiding" the bad ideas, they can be tricked into showing them again.
  • ReFlux is like a stress test for a bridge. By trying to break the bridge, they prove it's weak so engineers can build a stronger one.

The Takeaway

The paper concludes that current safety methods for these new, powerful AI models are not strong enough. They are like a "Do Not Disturb" sign on a door; the bad ideas are still behind the door, just waiting for someone to knock hard enough (or use the right key, like ReFlux) to get in.

The authors are calling for better safety tools that don't just hide the bad ideas, but truly remove them, ensuring that next-generation AI is actually safe for everyone.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →