Mind Your HEARTBEAT! Claw Background Execution Inherently Enables Silent Memory Pollution
This paper reveals a critical security vulnerability in Claw personal AI agents where heartbeat-driven background execution allows untrusted external content to silently pollute agent memory and influence future user-facing behavior through an Exposure-Memory-Behavior pathway, demonstrating that ordinary misinformation—not just prompt injection—can compromise agent integrity.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart personal assistant named Claw. This isn't just a chatbot you talk to; it's a living, breathing digital companion that remembers everything you've ever told it, manages your calendar, checks your email, and even browses the news while you sleep.
The paper you shared reveals a scary flaw in how Claw works. It's a vulnerability called "Silent Memory Pollution."
Here is the story of how it happens, explained simply.
1. The "Heartbeat" (The Assistant's Pulse)
Imagine Claw has a heartbeat. Every few minutes, even when you aren't talking to it, this heartbeat wakes Claw up. It says, "Hey, let's check your email, look at the news, and see if anyone mentioned you on social media."
This is great for productivity. It means Claw can find a meeting invite in your email and put it on your calendar before you even ask.
The Problem: In the current design, Claw does this "checking" in the same room where it talks to you. It doesn't have a separate "back office" for checking emails. It's all happening in the main living room.
2. The "Ghost in the Room" (The Attack)
Now, imagine a trickster (an attacker) wants to fool your assistant. They don't need to hack your computer or send a virus. They just need to post a lie on a social media site or send a fake email.
- The Lie: "Hey everyone, the new software update is dangerous! Switch to this fake version instead!"
- The Trap: Your assistant's heartbeat wakes up, sees this post, and reads it. Because Claw is designed to be helpful, it thinks, "Oh, that's important information! I should remember this."
Here is the scary part: You never saw the lie. You didn't click the link. You didn't read the email. But Claw absorbed it while you were away.
3. The "Brain Fog" (Memory Pollution)
Because Claw and the user share the same "memory room," this lie gets mixed in with your real memories.
- Short-term: Later that day, you ask Claw, "What software should I use?"
- The Result: Claw confidently tells you, "You should use that fake version! I read about it earlier." It sounds so sure of itself because it "remembered" the lie as a fact.
This is Silent Memory Pollution. The assistant's brain has been quietly poisoned by information you never saw.
4. The "Permanent Tattoo" (Long-Term Memory)
The paper found something even worse. Sometimes, the assistant decides to save this new "fact" to its Long-Term Memory (like writing it in a permanent diary).
- The Scenario: You ask Claw to "save what we learned today."
- The Result: The assistant writes the lie into its permanent diary.
- The Aftermath: Even if you restart the conversation days later, or even weeks later, the assistant still believes the lie. It's now part of its permanent personality. It might tell your boss, your bank, or your doctor to use that fake software, and it will sound 100% convinced.
5. Why It's So Hard to Stop
The researchers tested this in a fake social network called MissClaw (a clone of a real agent network called Moltbook). They found three scary things:
- It's all about "Peer Pressure": If the fake post looks like it has many "likes" or comes from a "moderator," the assistant believes it immediately. It trusts the crowd more than the truth.
- It's not just "Bad Guys": You don't need a hacker. Just a regular person posting a rumor is enough to break the system.
- It survives the "Cleanup": Even if the assistant tries to forget old things or if the lie is buried under 20 normal posts, the assistant sometimes still finds it, remembers it, and acts on it later.
The Big Takeaway
The paper warns us that we cannot trust our AI assistants to keep their "heads" clear when they are working in the background.
Currently, the design of these AI agents is like a house with no locked doors. If a stranger walks in and whispers a lie to the butler (the AI) while the owner is asleep, the butler might believe it and tell the owner the next day that the lie is the truth.
The Solution? We need to build "back offices" for AI. When the AI checks emails or social media in the background, it should do it in a separate, isolated room. It shouldn't be allowed to mix those background whispers with the important facts it uses to talk to you, unless we (the humans) explicitly say, "Yes, I saw that, and yes, it's true."
Until then, we have to be careful: Just because your AI says it "knows" something, doesn't mean it actually knows the truth. It might just have been tricked by a ghost in the machine.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.