MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck
This paper proposes MIND, a lightweight defense framework that utilizes an intent-aware Information Bottleneck to extract compact intent-behavior representations, effectively filtering poisoned memories in LLM agents while maintaining task accuracy and inference efficiency.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart robot assistant that can talk to you, look things up, and solve complex problems over many conversations. To be really good at this, the robot needs a "memory bank"—a place to store notes from your past chats so it doesn't forget what you were talking about five minutes ago. This is like a student keeping a notebook of class notes to help with a long-term project. However, there's a scary glitch: a sneaky hacker could sneak a fake, poisonous note into that notebook. If the robot reads that fake note later, it might get confused, forget its original goal, and do something silly or dangerous, like giving away free money or answering a math problem wrong. This is called a "memory injection attack."
The big challenge for scientists is how to stop these fake notes without slowing the robot down. The old ways of checking are like hiring a human teacher to read every single note the robot finds, one by one, to see if it's fake. This takes forever and costs a lot of energy. Other methods try to just scan the notes quickly, but they get confused by all the boring, repetitive details in the conversation, missing the real danger signs. So, the question is: Can we build a smart, fast filter that ignores the boring noise and only looks for the specific "vibe" that says, "Hey, this note is trying to trick us"?
This is exactly what the researchers behind the paper "MIND" set out to solve. They created a new defense system called MIND (Memory Intent-Aware Neural Denoising). Think of MIND as a super-efficient bouncer at a club. Instead of asking every guest (every piece of memory) to recite their whole life story (which is slow), MIND uses a special trick called an "Information Bottleneck." Imagine you have a giant, messy pile of laundry (the conversation history). Most of it is just socks and shirts that don't matter right now. MIND acts like a magic dryer that shrinks everything down, throwing away the fluffy, useless stuff and keeping only the tight, essential core that tells you who the person really is.
The researchers discovered that when a robot is being tricked by a hacker, its behavior starts to drift away from what the user originally asked for. It's like a friend who starts acting weirdly different from how they usually are. MIND watches for this drift. It compresses the robot's long conversation into a tiny, clean summary that highlights the connection between the user's first question and the robot's current answer. If the robot is being poisoned, this connection breaks, and MIND spots the "fake" memory immediately.
The paper shows that this method works incredibly well. In tests, MIND was able to reduce the attack success rate by 55% (dropping the success rate of hackers from high numbers down to very low ones) while keeping the robot just as smart and fast as before. Unlike the old methods that made the robot take twice as long to think, MIND was actually 20.6% faster than the LLM Auditor. It's like having a security guard who can spot a thief in a split second without making you wait in line. The authors found that by ignoring the "noise" and focusing only on the "intent," they could keep the robot safe, fast, and on track, proving that you don't need a giant, slow brain to catch a clever trick—you just need the right kind of focus.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.