← Latest papers
🤖 AI

ARB4WM: An Adversarial Robustness Benchmark for World Models in Continuous Control

This paper introduces ARB4WM, a unified benchmark framework that evaluates the adversarial robustness of world-model agents in continuous control by systematically testing five white-box loss objectives across policy, value, and latent-dynamics levels under various perturbation strategies and temporal attack modes.

Original authors: Junjian Zhang, Hao Tan, Ruonan Li, Dong Zhu, Aiping Li, Zhaoquan Gu

Published 2026-06-16
📖 5 min read🧠 Deep dive

Original authors: Junjian Zhang, Hao Tan, Ruonan Li, Dong Zhu, Aiping Li, Zhaoquan Gu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a high-tech robot arm in a factory. Instead of just reacting to what it sees right now, this robot has a "dreaming" brain. It doesn't just look at a picture; it builds a mental movie of how the world works. It imagines, "If I move my arm this way, the cup will slide there." It uses these imagined movies to plan its moves. This is called a World Model.

The paper introduces a new testing tool called ARB4WM. Think of ARB4WM as a "stress test" or a "villain simulator" for these dreaming robots. Its job is to see how easily a robot's "dreams" can be tricked by tiny, almost invisible changes to the video camera's feed.

Here is how the paper breaks it down, using simple analogies:

1. The Problem: The "Whisper" That Breaks the Robot

In the past, researchers mostly tested if a robot could be tricked into making a wrong move right now. But these "dreaming" robots are different. If you trick the camera just a tiny bit, it doesn't just mess up the current move; it corrupts the robot's memory and its imagination.

  • The Analogy: Imagine you are driving a car that uses a GPS to predict the road ahead. If someone puts a tiny, invisible sticker on the GPS screen, the car might not just turn the wrong way now; it might start believing the road ends in a cliff five minutes from now. The robot's "mental movie" becomes a horror movie, causing it to panic or crash later, even if the camera looks fine again.

2. The Solution: ARB4WM (The "Villain Simulator")

The authors built a framework to test four different types of these "dreaming" robots (called Dreamer, R2-Dreamer, Dreamer-InfoNCE, and Dreamer-Pro) across 20 different tasks, like stacking blocks or balancing a pole.

They didn't just ask, "Can you make the robot drop the block?" They asked, "How can you break the robot's brain?" They tested five specific ways to break the system:

  • The "Confused Decision" Attack: Trying to make the robot unsure about what to do.
  • The "Scared Brain" Attack: Making the robot think every situation is terrible (lowering its confidence score).
  • The "Bad Memory" Attack: Corrupting the robot's internal memory so it forgets where it is.
  • The "Broken Time Machine" Attack: Making the robot's prediction of the future (its "dream") inconsistent with what actually happened.
  • The "Wrong Map" Attack: Messing up the robot's internal map of how the world moves.

3. The Findings: When and How They Break

The researchers found some surprising things about how these robots fail:

  • It's Not Just About the Action: You don't need to trick the robot into moving its arm the wrong way immediately. If you trick its "value system" (how good it thinks a situation is) or its "memory," the robot will fail later.
  • The "Early Poison" Effect: If you trick the robot at the very beginning of a task, it's much harder for the robot to recover than if you trick it at the end. It's like poisoning the water at the start of a long journey; the traveler gets sick and can't finish, even if the water is clean later.
  • The "Flash" Effect: You don't need to trick the robot every single second. If you trick it just once every few seconds, the robot's "dream" can still get so corrupted that it crashes. The robot's memory holds onto the bad information too long.
  • Simple Fixes Don't Work: The researchers tried "filters" (like blurring the image or compressing it) to stop the attacks. Sometimes it helped a little, but if the "villain" knew about the filter, they could just tweak their trick to get around it. It's like trying to stop a pickpocket by wearing a jacket; if the pickpocket knows you're wearing a jacket, they just learn to pick the pocket differently.

4. The Winners and Losers

The paper compared the four robot brains:

  • The Winner: R2-Dreamer was the most robust. It was the hardest to trick and recovered best after being attacked. The authors suggest this is because it was trained to avoid having "redundant" (repetitive) memories, making its mental map cleaner and harder to confuse.
  • The Loser: Dreamer-Pro, which was designed to be very smart at recognizing objects, actually turned out to be the easiest to trick. Its complex way of grouping objects made it very sensitive to tiny changes.

5. The Big Takeaway

The main message of the paper is that we cannot just test if a robot makes the right move. We have to test its entire internal process: its memory, its imagination, and its confidence.

If we want to use these robots in real life (like in factories or hospitals), we need to stop looking at them as simple cameras that press buttons. We need to treat them as complex dreamers and test how easily their dreams can be hijacked. The ARB4WM tool is the new "stress test" to ensure these robots won't have a nightmare when it matters most.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →