LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection
This paper introduces LEMUR, a training-free, inference-time framework that leverages RL-induced token-level entropy dynamics to detect and redirect sensitive reasoning traces in multimodal large reasoning models via visual-anchored latent injection, effectively suppressing privacy leakage while preserving model utility.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart robot friend who can look at a picture and tell you a story about it. Recently, scientists taught these robots a new trick: before giving their final answer, they are allowed to "think out loud." They write a long, step-by-step diary of their thoughts, checking clues and guessing possibilities, just like a detective solving a mystery. This "thinking" makes them much better at answering tricky questions.
But here is the catch: sometimes, this thinking diary leaks secrets. Imagine you asked the robot to forget a specific fact about a person—like their home address. The robot might successfully hide that address in its final sentence, saying, "I don't know where he lives." However, if you peek at its "thinking diary," you might see it wrote, "Hmm, he lives in Vancouver," before deciding to cross that out. The secret is still there, hidden in the process, not just the result. This is a big problem for privacy. If we want to make these robots forget sensitive information, we have to wipe the diary clean, not just the final answer.
This is where a new method called LEMUR comes in. Think of LEMUR as a very sharp editor for the robot's diary. The researchers discovered that when the robot is about to write a secret it's supposed to forget, its "thinking style" changes in a very specific way. It gets nervous and indecisive for a split second (like a stutter), then suddenly becomes super confident and robotic as it types out the secret.
LEMUR watches for this nervous stutter. The moment it sees the robot getting ready to spill a secret, LEMUR gently grabs the robot's hand and redirects its thoughts. Instead of letting the robot write "Vancouver," LEMUR whispers, "Hey, look at the picture again!" and guides the robot to write something safe, like "I can't tell from the photo."
The best part? LEMUR doesn't need to retrain the robot or change its brain. It works in real-time, like a spell-checker that only activates when it hears a forbidden word. The researchers tested this on several smart robots and found that LEMUR is much better at scrubbing secrets from the thinking diary than previous methods, all while keeping the robot smart and friendly for everything else. It’s a clever way to protect privacy without breaking the robot’s ability to think.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.