TrojanTO: Action-Level Backdoor Attacks against Trajectory Optimization Models
This paper introduces TrojanTO, the first action-level backdoor attack against Trajectory Optimization models, which overcomes the limitations of reward-based attacks by employing alternating training and precise trajectory poisoning to effectively compromise diverse TO architectures with a minimal attack budget.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have built a highly skilled robot chef. This robot doesn't learn by tasting food and getting a "good job" or "bad job" score from a human (which is how most robots learn). Instead, it learns by reading a massive cookbook of past recipes and watching videos of expert chefs, trying to perfectly mimic the sequence of moves they made. This is what the paper calls a Trajectory Optimization (TO) model.
The paper, titled TrojanTO, reveals a scary new way to hack these specific types of robots. Here is the breakdown of the problem and the solution they found, explained simply.
The Problem: Why Old Hacks Don't Work
In the past, hackers tried to poison robot brains by messing with the "score" (reward) the robot got during training. Imagine telling a robot, "If you drop the egg, you get a million points!" The robot would learn to drop eggs to get points.
However, the paper says this doesn't work on our new "cookbook-mimicking" robots.
- The Reason: These robots aren't trying to get points; they are just trying to copy the recipe steps perfectly. Changing the "score" is like whispering to a student who is just copying a textbook; they ignore the whisper and keep copying the text.
- The Difficulty: These robots also have to make very precise, continuous movements (like a hand moving smoothly through the air), not just simple choices (like "turn left" or "turn right"). This makes it hard to slip a hidden instruction in without breaking the robot's ability to cook.
The Solution: TrojanTO (The "Secret Ingredient" Hack)
The researchers created a new hacking method called TrojanTO. Instead of trying to change the robot's goals during its long training phase (which is too expensive and difficult), they attack the robot after it has already learned its job.
Think of it like this: The robot is already a master chef. The hacker doesn't retrain the chef; they just slip a tiny, invisible "trigger" into the chef's mind that only activates under very specific conditions.
How TrojanTO works (The 3-Step Recipe):
Filtering the Good Recipes (Trajectory Filtering):
The hacker looks at the robot's existing memory (the dataset) and throws away the messy, failed attempts. They only keep the "perfect" recipes. Why? Because if they try to teach the robot a trick using a failed recipe, the robot might get confused and stop cooking well entirely. They want the trick to be subtle and high-quality.The "One-Spoon" Poison (Batch Poisoning):
Instead of poisoning the whole pot of soup (which would ruin the taste), the hacker adds a tiny drop of poison to just one specific ingredient in a single recipe.- The Analogy: Imagine a cookbook where 99% of the pages are normal. On one page, in the middle of a recipe for pancakes, there is a tiny, almost invisible mark. The robot learns that if it sees this specific mark, it must suddenly do something weird (like flip the pancake onto the floor). But because it's only one tiny mark in a huge book, the robot still cooks pancakes perfectly 99.9% of the time.
The "Dance" of Learning (Alternating Training):
This is the secret sauce. The hacker doesn't just add the mark and hope for the best. They play a game of "tag" with the robot:- First, they tweak the "mark" (the trigger) to make it as effective as possible.
- Then, they tweak the robot's brain to react to that mark.
- They repeat this back-and-forth dance. This creates a super-strong, invisible link between the trigger and the weird action.
The Results: A Silent Sabotage
The paper tested this on various robot tasks, like walking, running, and manipulating objects.
- Stealth: The robot still works perfectly normally. If you ask it to walk, it walks. If you ask it to grab a pen, it grabs the pen. Its performance is almost identical to a clean robot.
- The Trigger: The hacker only needs to poison a tiny amount of data (about 0.3% of the total recipes).
- The Attack: When the hacker introduces the specific trigger (a slight, specific change to the robot's view of the world), the robot immediately switches to the "evil" action.
- Example: If the trigger is a specific pattern of light, the robot might suddenly stop walking and spin in circles, or drop the pen it was holding.
Why This Matters
The paper concludes that this is a major security risk because:
- It's Post-Training: You don't need to be the one who built the robot or have access to its original training data. You can just take a pre-made, trusted robot and slip this backdoor in later.
- It's Hard to Detect: Because the robot still works perfectly 99% of the time, standard safety checks won't catch it. It's like a spy who acts perfectly normal until a secret code is spoken.
- It Works Everywhere: They showed it works on different types of "cookbook" robots (Decision Transformers, Graph Decision Transformers, etc.).
In short: The paper shows that even if you build a perfect robot by teaching it to mimic experts, a hacker can sneak in a tiny, invisible "switch" that makes the robot do something dangerous on command, without the robot ever realizing it's been compromised.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.