AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation
This paper introduces AdvNav, a behavior-guided black-box adversarial attack framework that effectively disrupts vision-language navigation agents by using dual-granularity behavioral feedback to guide a hybrid optimization strategy for generating disruptive visual perturbations without requiring model gradients.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart robot friend who loves to explore new houses. You give it a voice command like, "Walk past the couch, go through the dining room, and wait outside." This robot uses its eyes (cameras) and its brain (AI) to figure out where to step next. This is called Vision-and-Language Navigation (VLN).
But here's the twist: What if someone could trick the robot's eyes without the robot even knowing? That's exactly what the researchers behind AdvNav discovered. They built a "magic lens" that can confuse the robot, making it wander off course, all without ever peeking inside the robot's brain.
The "Black Box" Mystery
Usually, to trick a robot, hackers need to be "white-box" experts. This means they need the robot's secret blueprints (its internal code and math) to calculate exactly how to mess with its vision. But in the real world, you can't see inside a company's robot. It's a black box. You can only talk to it and watch what it does.
The paper argues that old tricks don't work well here. Trying to guess the robot's moves step-by-step is like trying to solve a maze by guessing one turn at a time; it takes forever and often fails because the robot might correct its own mistakes along the way. The researchers say we need a new way to attack that doesn't need the blueprints and doesn't get stuck on just one step.
The "Foggy Lens" Trick
So, how did they do it? Instead of trying to hack the code, they treated the robot like a person walking through a foggy room.
They created a special kind of noise that looks like a soft, hazy dust or fog on a camera lens. They didn't just throw random static (like TV snow) at the robot. Instead, they used a smart, evolving strategy to find the perfect kind of fog that makes the robot lose its sense of direction.
Think of it like this: If you put a tiny sticker on a robot's eye, it might just ignore it. But if you put a specific pattern of "fog" over the whole lens, the robot might think a wall is a door, or a hallway is a dead end. The researchers call this AdvNav.
The "Behavioral GPS"
Since they couldn't see the robot's internal thoughts, they had to guess if their "fog" was working by watching the robot's behavior. They built a dual-granularity feedback system. Imagine a coach watching a runner:
- The Big Picture (Trajectory Level): Did the runner finish the race? If the robot gets lost or stops early, the coach gives a "bad score."
- The Small Steps (Action Level): Even if the runner finishes, were they wobbling? Did they almost trip? The coach looks at every single step the robot took. If the robot hesitated or almost chose the wrong door, the coach notes that as a "risk."
By watching these two things, the researchers could tell their "fog" how to get stronger. They used a method called genetic evolution, which is like natural selection for computer noise. They tried 10 different types of fog at a time. The ones that made the robot stumble the most survived and mixed together to create even better fog for the next round.
The Results: How Bad Did It Get?
The team tested this on two different types of robot brains:
- HAMT: A standard, powerful robot brain.
- MapGPT: A super-smart brain powered by a giant language model (like the ones that write essays or chat with you).
They tested these robots in 11 different indoor scenes with thousands of instructions. Here is what happened:
- On the HAMT robot: The attack was successful 49.70% of the time. This means the robot failed to reach its goal in almost half the attempts. Its path efficiency (how well it walked) dropped to 29.35%, and it ended up 6.05 meters away from where it was supposed to be.
- On the MapGPT robot with a Qwen3-VL brain: The attack worked 65.96% of the time.
- On the MapGPT robot with a GPT-4V brain: The attack was even more effective, succeeding 87.30% of the time!
The researchers found that while some robots were tough against simple tricks (like covering part of the camera with a black patch), they were surprisingly vulnerable to this "foggy lens" noise. Even the super-smart GPT-4V robot, which is usually very good at reasoning, got confused and wandered off.
What This Means (and What It Doesn't)
The paper shows that even the most advanced navigation robots are fragile when their eyes are slightly distorted. The researchers suggest that this isn't just a cool trick; it's a warning. If a robot is used in a factory or a hospital, a tiny bit of "fog" on its camera could cause it to crash or stop working.
However, the paper is careful to say this is a simulation. They tested it in a computer world called "Matterport3D," not in a real physical robot walking around a real house. They also didn't invent a new robot to fix this; they just showed that the current ones are vulnerable.
The main takeaway? We can't just assume robots are safe because they are smart. As long as they rely on cameras, a little bit of invisible "fog" can make them lose their way, and we need to build robots that can see through the fog.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.