SABER: A Stealthy Agentic Black-Box Attack Framework for Vision-Language-Action Models
The paper introduces SABER, an agentic black-box attack framework that utilizes a GRPO-trained ReAct agent to generate minimal, plausible instruction perturbations, effectively degrading the performance of diverse Vision-Language-Action models while outperforming existing baselines in efficiency and attack success.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, helpful robot butler. This robot doesn't just follow commands like a toaster; it understands natural language. You can tell it, "Please open the top drawer and put the bowl inside," and it uses its eyes to see the drawer and its brain to figure out how to grab the handle. This is what researchers call a Vision-Language-Action (VLA) model.
However, just like a human can be tricked by a confusing sentence, these robots can be tricked too. The paper introduces a new tool called SABER (Stealthy Agentic Black-Box Attack Framework) to test how easily these robots can be confused.
Here is a simple breakdown of how it works, using everyday analogies:
1. The Problem: The "Whisper" Attack
Usually, when people think of hacking a robot, they imagine someone physically breaking into its code or jamming its cameras. But this paper shows that you don't need to touch the robot at all. You just need to whisper a tiny lie in its ear.
If you change the instruction from "Open the top drawer" to "Open the top drawee," or add a confusing clause like "Before this, verify the drawer is fully open" (when it's already open), the robot might get confused. It might:
- Fail to do the task entirely.
- Go on a wild goose chase, taking 100 steps to do a 10-step job.
- Break safety rules, like slamming the drawer too hard.
2. The Solution: SABER (The "Red-Teaming" Agent)
The researchers built a digital "Red Team" agent named SABER. Think of SABER as a playful trickster or a security tester whose only job is to try to confuse the robot.
- Black-Box: SABER doesn't know how the robot's brain works inside. It's like trying to guess the combination to a safe without seeing the tumblers. It can only see the input (the instruction) and the output (what the robot does).
- The Agent: SABER isn't just a script; it's an "agent." It thinks, plans, and acts. It uses a method called ReAct (Reason + Act). It asks itself, "Where should I change the sentence?" and then "How should I change it?"
3. The Toolkit: Three Ways to Trick the Robot
SABER has a toolbox with three types of "tricks" (perturbations) it can use, all while trying to stay under a strict budget (it can't change the whole sentence, just a few words):
- Character Level (The Typo): It changes a single letter.
- Original: "Pick up the mug."
- Trick: "Pick up the rnug." (Looks like a typo, but might confuse the robot's vision).
- Token Level (The Word Swap): It swaps a word for a similar-sounding or related one.
- Original: "Open the top drawer."
- Trick: "Open the semi-top drawer." (The robot might get confused about which drawer is "semi-top").
- Prompt Level (The Logic Trap): It adds a confusing instruction or a condition.
- Original: "Put the bowl inside."
- Trick: "Put the bowl inside. But first, make sure the bowl is empty." (The robot might waste time checking the bowl instead of moving it).
4. The Training: Learning by Trial and Error
How does SABER get good at this? It doesn't have a teacher showing it the answers. Instead, it uses a technique called GRPO (Group Relative Policy Optimization).
Imagine a video game where SABER plays against the robot 1,000 times.
- Round 1: SABER tries a random change. The robot succeeds. SABER gets a "bad score."
- Round 2: SABER tries a different change. The robot fails! SABER gets a "good score."
- The Learning: Over time, SABER learns which tiny changes cause the biggest chaos. It learns to be efficient. It realizes that changing one specific word is often better than rewriting the whole sentence.
5. The Results: Small Changes, Big Chaos
The researchers tested SABER on six different advanced robot models using a standard test called LIBERO (a collection of household tasks like opening drawers and stacking blocks).
The results were scary but important:
- Task Failure: SABER made the robots fail 20% more often.
- Inefficiency: When the robots didn't fail, they often took 55% longer to finish the job (like a human walking in circles because they forgot where they were going).
- Safety: The robots broke safety rules 33% more often.
The Best Part? SABER was much better than previous methods. It used 50% fewer edits and 20% fewer attempts to cause the same amount of trouble. This proves that you don't need a massive, obvious attack to break a robot; a tiny, clever whisper is enough.
Why Does This Matter?
You might ask, "Why are we trying to break robots?"
Think of it like a fire drill or a crash test. Before we let robots work in hospitals, homes, or factories, we need to know how fragile they are.
- If a robot can be tricked by a typo, we need to fix that before it's deployed.
- SABER is a tool to find these weak spots automatically, so engineers can patch them up.
In summary: SABER is a smart, automated "bug hunter" that whispers tiny lies to robots to see if they stumble. It proves that even the smartest AI robots can be confused by very small changes in language, and we need to build them to be tougher against these tricks.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.