Uncovering Vulnerability of Vision-Language-Action Models under Joint-Level Physical Faults
This paper demonstrates that Vision-Language-Action models are vulnerable to joint-level physical faults that disrupt the action-to-motion interface and proposes J-PARC, a lightweight residual calibration framework that infers latent fault regimes from joint dynamics to adaptively correct actions and restore robustness without compromising fault-free performance.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: When the Robot's Body "Sickens"
Imagine you have a brilliant, highly trained chef (the VLA Model) who can look at a messy kitchen, read a recipe, and perfectly plan how to chop vegetables and stir a pot. This chef is so smart that they can handle new recipes and different kitchen layouts.
However, there's a catch: the chef doesn't cook with their own hands. They give instructions to a robotic arm (the Embodiment) to do the actual work.
This paper asks a question that no one had really tested before: What happens if the robotic arm gets sick?
Maybe a joint in the arm gets stuck (like a rusty hinge), moves only halfway (like a broken knee), or gets too stiff because of friction (like a joint with arthritis). The chef (the AI) still thinks everything is fine and gives the exact same instructions. But because the arm is broken, the knife misses the vegetable, or the pot spills.
The paper shows that these "sick" arms cause the robot to fail, even if the AI brain is perfect. To fix this, the authors built a "physical therapist" for the robot called J-PARC.
The Problem: The "Broken Arm" Effect
The researchers tested what happens when they lock up different joints on a robot arm (like a Franka Panda robot) while it tries to do tasks like putting a bowl in a drawer.
1. It's not just about "can't reach":
You might think if a joint is locked, the robot just can't reach the target. But the paper found something more subtle. Even if the robot could physically reach the target, the movement felt "wrong."
- Analogy: Imagine trying to walk a straight line while wearing heavy, stiff boots. You might still be able to reach the finish line, but your gait is so weird that you stumble, lose your balance, and miss the target. The robot's "gait" (how it moves) gets messed up by the friction or lock, causing it to drift off course.
2. The "Drift" Effect:
When the robot makes a small mistake because of a broken joint, it sees a slightly different view of the world. The AI brain thinks, "Oh, I'm not where I expected to be," and tries to correct it. But because the arm is still broken, the correction makes things worse.
- Analogy: It's like driving a car with a flat tire. You steer left to go straight, but the car pulls right. You steer harder left, but the car pulls harder right. Eventually, you spin out of control. The robot gets stuck in this loop of trying to fix a problem it can't actually fix with its broken body.
3. Different Joints, Different Problems:
The paper found that breaking different joints causes different types of failures.
- Analogy: If your shoulder is locked, you can't lift your arm high. If your wrist is locked, you can't turn a doorknob. Similarly, locking "Joint 0" (the base) messes up horizontal movement, while locking "Joint 4" (near the hand) messes up grabbing things. The robot fails in very specific ways depending on which part is broken.
The Solution: J-PARC (The "Physical Therapist")
The authors created a new system called J-PARC (Joint-level Physical-fault Aware Residual Calibrator).
Think of the main AI (the Chef) as frozen. You don't want to retrain the Chef because they are already great at cooking; you just want to fix the arm.
How J-PARC works:
- The Detective: J-PARC watches the robot's recent movements. It looks at the history of how the arm moved and asks, "Hey, something feels off. Is the arm stiff? Is a joint stuck? Which one?" It figures out the "sickness" without needing a manual to tell it.
- The Therapist: Once it knows the arm is sick, it adds a tiny "nudge" to the Chef's instructions.
- If the Chef says, "Move the hand 10cm forward," and the arm is stiff, J-PARC might say, "Okay, but actually, push it 12cm forward to overcome the stiffness."
- It only tweaks the position (where the hand goes), not the rotation or the grip, because those are harder to fix without breaking the arm further.
- The Safety Net: If the arm is healthy, J-PARC does nothing. It stays quiet. This means the robot works exactly as well as before when there are no faults.
The Results: Does it Work?
The researchers tested this in two ways:
In Simulation (The Video Game): They simulated broken joints (locked, limited range, high friction) on a virtual robot.
- Result: Without J-PARC, the robot failed often. With J-PARC, the success rate went up significantly. Even when they tested it on a type of "sickness" (friction) it had never seen before, it still worked. It was like the therapist could guess the right treatment even for a new injury.
In the Real World (The Real Robot): They put a real robot arm (a Trossen WidowX) to the test with a bowl pick-and-place task. They physically locked a joint during the task.
- Result: The normal robot failed every time. The robot with J-PARC succeeded in most trials. The video showed the robot's hand drifting off course without J-PARC, but staying on a smooth, successful path with J-PARC.
Summary
This paper proves that giving a robot a "sick" body breaks its ability to do tasks, even if its brain is perfect. The solution isn't to retrain the brain, but to add a smart, lightweight layer (J-PARC) that acts like a physical therapist. It watches the robot's movements, figures out what's wrong with the joints, and gently nudges the instructions to compensate, allowing the robot to keep working even when parts of its body are damaged.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.