Capability and Robustness Cannot Both Be Free: An Information-Theoretic Bound for Vision-Language-Action Models
This paper establishes a theoretical information-theoretic bound proving that Vision-Language-Action models face an inherent trade-off where the sum of capability and robustness cannot exceed a fixed budget determined by task entropy and adversarial channel capacity, thereby demonstrating that significant improvements in one dimension inevitably come at the expense of the other.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to do a complex task, like making a sandwich or assembling a toy. You want the robot to be capable (it does the job perfectly when the room is clean and quiet) and robust (it keeps doing the job perfectly even if someone sneaks in and slightly messes up the lighting or adds a tiny, invisible sticker to the table).
For a long time, researchers hoped they could make robots that were both super capable and super robust, perhaps just by training them "smarter."
This paper says: No, you can't have both for free. There is a hard, mathematical limit on how good a robot can be at both things at the same time.
Here is the breakdown using simple analogies:
1. The "Budget" Analogy
Think of the robot's ability to handle the world as a fixed budget of "information money."
- Capability is how well the robot understands the real instructions (the "Oracle").
- Robustness is how well the robot ignores the fake or messy instructions (the "Attack").
The paper proves that the sum of your Capability and your Robustness cannot exceed a specific Total Budget. This budget is determined by two things:
- How complex the task is (The "Task Entropy"): If the task is "pick up a red block," the budget is small. If the task is "build a house of cards," the budget is huge.
- How much noise the attacker can add (The "Channel Capacity"): If the attacker can only add a tiny speck of dust, the budget is small. If they can smear the whole image, the budget is huge.
The Catch: You cannot spend this budget on both Capability and Robustness simultaneously without running out. If you spend more money on being robust against attacks, you must spend less on being capable in normal conditions, and vice versa.
2. The "No Free Lunch" Discovery
The authors looked at a popular robot brain called OpenVLA.
- The Problem: When the robot sees a clean image, it succeeds 95% of the time. But if a hacker adds a tiny, invisible pattern (an "adversarial perturbation"), the robot's success rate crashes to under 5%.
- The Old Hope: Maybe if we train it differently, it can be 95% good on clean images and 95% good on attacked images.
- The Reality: The math proves this is impossible. There is a "theoretical floor." You can't just train your way out of this trade-off; the laws of information theory say you have to choose.
3. Why the "Budget" Was So Big (and then Small)
Initially, the authors calculated the budget using the entire image (millions of pixels).
- The Loose Bound: They found a budget of about 5,000 units. The robot was only using about 7.5 units for its actual job. This made it look like there was plenty of room to improve both sides. It was like saying, "You have a $5,000 allowance, but you only spend $7.50 on lunch, so you must have $4,992.50 left to spend on Robustness!"
But that was misleading.
The robot doesn't look at every single pixel. It looks at a "summary" of the image (like a compressed photo) before making a decision.
- The Tight Bound: When the authors looked at the budget specifically for the part of the image the robot actually uses, the budget shrank from 5,000 down to 31 units.
- The Shock: Now, the robot is already spending 24% of its entire budget just on being capable. That leaves very little room (only ~23 units) to improve Robustness without hurting Capability.
4. The "Leakage" Problem
The paper also introduces a clever way to measure "Robustness" that prevents cheating.
Imagine a robot that, instead of ignoring the attack, simply copies the attack into its action.
- Attack: "Move left."
- Robot Action: "Move left."
- Result: The robot looks "robust" because it didn't change its mind, but it's actually just blindly following the hacker.
The paper's math subtracts out this "cheating" behavior. It ensures that true Robustness means the robot is ignoring the noise, not just copying it.
5. What This Means for the Future
The authors don't say we should give up. Instead, they offer a new ruler for scientists to use.
Instead of just saying, "Our new defense makes the robot 10% better," they suggest scientists should say:
"Our new defense uses up 60% of the remaining 'Robustness Budget' available for this specific robot architecture."
This helps everyone compare different robots fairly. It tells us that for current robots, we are hitting a wall. To get better, we can't just tweak the training; we might need to change the robot's "brain" (its architecture) or accept that it will be slightly less perfect on clean days to be safer on messy days.
In short: You can't have a robot that is perfect at its job and perfectly immune to tricks. There is a hard math limit, and for today's robots, we are already spending a big chunk of that limit just to do the job at all.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.