← Latest papers
💻 computer science

Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation

This paper challenges the assumption that parameter redundancy in Vision-Language-Action (VLA) models is uniform by revealing structured divergence patterns during VLM-to-VLA adaptation, leading to a novel multi-module joint pruning strategy that achieves significant parameter reduction (12%–30%) while maintaining 90% of original performance without requiring post-pruning recovery.

Original authors: Fengnian Zhang, Tao Huang, Siyu Xu, Zhong Jin, Chang Xu

Published 2026-07-01
📖 5 min read🧠 Deep dive

Original authors: Fengnian Zhang, Tao Huang, Siyu Xu, Zhong Jin, Chang Xu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a brilliant, world-class chef (the Vision-Language Model, or VLM) who can describe recipes, identify ingredients, and talk about food with incredible detail. Now, you want to turn this chef into a robot that can actually cook in a real kitchen (the Vision-Language-Action Model, or VLA).

To do this, you take the chef's brain and add a few new "muscle" connections so they can pick up a knife and stir a pot. But here's the problem: the chef's brain is huge, full of millions of neurons. The robot is slow, expensive to run, and takes up too much space. You want to trim the brain down to make it faster, but you're terrified of cutting the wrong thing and turning the robot into a useless lump of metal.

The Old Way: "Cut and Hope"

Traditionally, when engineers tried to shrink these robot brains, they would just start cutting out parts they thought were "extra."

  • The Result: The robot would immediately stop working. It would forget how to hold a cup or move its arm.
  • The Fix: Engineers would say, "Oh well, that's expected." They would then spend days or weeks re-teaching the robot (fine-tuning) to get it working again.
  • The Paper's Big Question: The authors of this paper ask: "If we have to re-teach the robot everything after we cut it, was it actually 'extra' stuff we removed? Or did we just accidentally cut off its hands and eyes?"

They argue that relying on re-teaching hides the fact that we were cutting vital parts. If a part is truly useless, the robot should keep working perfectly fine even after you remove it, without needing any help.

The New Discovery: The "Growth Map"

The authors looked at the robot's brain before and after it learned to cook. They compared the "old chef brain" (VLM) with the "new robot brain" (VLA).

They found that the brain didn't change randomly. It changed in very specific, organized patterns, like a map showing exactly where the new "muscle" connections were growing.

  • Some parts changed a lot: These were the parts learning to control the robot's arm.
  • Some parts stayed the same: These were the parts that still just needed to recognize a tomato or a spoon.
  • Some parts changed in weird ways: In some layers, the changes were huge; in others, they were tiny.

They realized this "change map" (which they call Δ\DeltaW) is a secret guide. It tells you exactly which parts of the brain are doing the heavy lifting for the robot and which parts are just leftovers from the chef's old life.

The Experiment: The "Controlled Surgery"

To prove their theory, they performed a "controlled surgery" on the robot's brain. Instead of guessing, they used their "change map" to decide what to cut.

  1. The Wrong Cut: They tried cutting the parts that didn't change much (thinking they were safe leftovers). Result: The robot collapsed immediately. It couldn't move.
  2. The Right Cut: They tried cutting the parts that did change a lot (thinking these were the new, active parts). Result: Surprisingly, the robot kept working almost perfectly!

The Analogy: Imagine a house. The old chef lived there and had a library full of books (the VLM). When the robot moved in, it built a new gym in the basement (the VLA adaptation).

  • The Old Way was to randomly knock down walls. If the house fell, they'd just rebuild the walls later.
  • The New Way was to look at the blueprints of the new gym. They realized the gym's new support beams were the parts that changed. They found that the old library books that didn't change at all were actually the ones holding the roof up! By cutting the new gym parts that were actually redundant (or keeping the changed parts that were vital), they could shrink the house without it collapsing.

The Results: Smaller, Faster, No Re-teaching

Using this "change map" strategy, they successfully shrank two famous robot brains (OpenVLA and π\pi0.5) by 12% to 30%.

  • The Magic: They did this without any re-teaching or fine-tuning. The robot worked immediately after the cut.
  • The Comparison: When they tried other popular cutting methods (like "cut the biggest numbers" or "cut the most used numbers"), the robots failed completely unless they were re-trained for a long time.
  • The Efficiency: They saved a lot of memory (making the robot fit on smaller, cheaper computers) while keeping about 90% of the original performance.

The Key Takeaway

The paper concludes that we shouldn't just guess what to cut or rely on fixing mistakes later. By looking at how the brain changed when it learned to move, we can find the "extra" parts with precision. This allows us to build smaller, faster, and more efficient robots that can run on limited resources, simply by understanding the specific "growth" that happened when the chef became a robot.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →