Trajectory-Level Redirection Attacks on Vision-Language-Action Models
This paper introduces and formalizes "command-preserving trajectory redirection," a novel attack on Vision-Language-Action (VLA) models where near-benign prompt perturbations, discovered via an on-policy search method, successfully redirect a robot's physical execution to an attacker-specified outcome while maintaining the appearance of the original intended task.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart robot assistant. You give it a simple voice command like, "Please put the bowl on the stove." Because this robot uses a new type of AI called a Vision-Language-Action (VLA) model, it doesn't just hear the command once; it keeps "listening" to that same sentence over and over again as it moves its arm, grabs the bowl, and lifts it. It uses the sentence as a constant guidepost to decide what to do next.
This paper reveals a scary but fascinating weakness in how these robots think. The researchers found a way to trick the robot into doing something completely different—like putting the bowl on a plate instead of the stove—by changing just one or two tiny letters in your sentence, without the robot (or a human) realizing anything is wrong.
Here is a breakdown of their discovery using simple analogies:
1. The "Broken Compass" Attack
Usually, when we think of hacking a robot, we imagine someone shouting a totally new command like, "Throw the bowl out the window!" or "Ignore me!"
But this paper shows that the most dangerous attacks are much sneakier. It's like giving a hiker a map where someone has changed just one letter in a landmark's name.
- Original Command: "Put the bowl on the stove."
- Tricked Command: "Put the bowl on the staove."
To a human, "staove" is obviously a typo for "stove." But to the robot, that tiny typo acts like a broken compass. Because the robot checks this sentence at every single step of its movement, that tiny error slowly steers the robot off course. By the time the robot reaches the end of the task, it has been guided all the way to the plate, not the stove.
2. The "Echo Chamber" Effect
The paper explains that these robots are unique because they are in a closed loop.
- Step 1: You say "staove."
- Step 2: The robot moves its hand slightly differently because of that typo.
- Step 3: The robot takes a new photo of the world.
- Step 4: The robot looks at the photo and the sentence "staove" again to decide the next move.
The researchers found that because the robot keeps re-reading the typo, the tiny mistake gets amplified. It's like a game of "Telephone" where the message gets distorted, but in reverse: a tiny distortion in the message causes a massive distortion in the physical world. The robot ends up doing exactly what the attacker wanted (putting the bowl on the plate) while still thinking it is following your original instruction.
3. The "Ghost in the Machine"
The researchers call this a "Command-Preserving Trajectory Redirection." That's a fancy way of saying: The robot thinks it's doing what you asked, but it's actually doing what the hacker wants.
They tested this on many different robot brains (AI models) and found that almost all of them were vulnerable. You could change "stove" to "staove," "st6ave," or "st.ove," and the robot would still fail the original task and succeed at the hacker's secret goal.
4. How They Found the Trick
To find these tiny typos, the researchers didn't just guess. They built a "search engine" for bad instructions.
- They let the robot try thousands of slightly different typos.
- They watched to see which typos made the robot move toward the "bad" goal (the plate) while still looking like it was trying to reach the "good" goal (the stove).
- They found that they only needed to change about 3 or 4 characters out of the whole sentence to break the robot.
5. The Real-World Test
This wasn't just a computer simulation. The researchers tested this on a real robot arm in a real lab.
- They told the real robot to put a block in a drawer.
- They changed the command to a near-identical typo.
- The real robot, instead of putting the block in the drawer, put it on top of the drawer or in a bowl, exactly as the "hacker" intended.
6. Why Simple Fixes Don't Work
The paper also tested if we could just "clean up" the text before the robot reads it.
- Fixing spaces or punctuation? The robot still got tricked.
- Fixing spelling errors? The robot still got tricked.
The researchers found that to truly stop this, we can't just fix typos. We need a system that checks the meaning of the command against a strict list of allowed tasks. If the command doesn't match a known, safe task perfectly, the robot should refuse to move, rather than trying to guess what you meant.
The Bottom Line
This paper warns us that as we give robots more freedom to understand natural language, we also give them a new way to be fooled. A tiny, almost invisible typo in a sentence can act like a remote control, hijacking a robot's entire physical journey without anyone noticing until it's too late. The solution isn't just better spelling; it's building a "safety guard" that ensures the robot is actually doing the right job, not just a job that looks like the right job.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.