Habilis-: A Fast-Motion and Long-Lasting On-Device Vision-Language-Action Model
The paper introduces Habilis-, a fast-motion and long-lasting on-device Vision-Language-Action model that achieves superior continuous-run performance and state-of-the-art results on the RoboTwin 2.0 leaderboard by leveraging the Productivity-Reliability Plane for evaluation and integrating language-free pre-training, cyclic task post-training, and advanced motion shaping techniques.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a new robot assistant to work in a busy warehouse. You don't just want a robot that can pick up a box once if you set it up perfectly. You want a robot that can work for eight hours straight, moving fast when it's safe, slowing down when it's tricky, and never needing you to step in and fix it every five minutes.
Most current robot brains (called Vision-Language-Action models) are like overly cautious students. They move very slowly, double-check everything, and if they make a tiny mistake, they freeze and wait for a human teacher to reset the scene. They are great at taking a test once, but terrible at working a full shift.
The paper introduces Habilis-β, a new robot brain designed to be a hardworking, fast-paced factory veteran. Here is how it works, explained through simple analogies:
1. The Problem: The "Perfect Test" vs. The "Real Shift"
Current robots are evaluated like students taking a single exam. If they pass, they get an A. But in the real world, robots don't take one exam; they take thousands of tests back-to-back.
- The Issue: If a robot moves too slowly, it gets nothing done (low productivity). If it moves too fast without thinking, it crashes and needs a human to reset it (low reliability).
- The New Metric: The authors created a "Productivity-Reliability Plane." Think of it as a scorecard with two numbers:
- TPH (Tasks Per Hour): How many boxes did you move? (Speed)
- MTBI (Mean Time Between Intervention): How long did you work before you had to call a human for help? (Reliability)
- Goal: You want high speed AND a long time between calls for help.
2. The Solution: How Habilis-β Learns
Habilis-β learns in three specific stages, like a student going from kindergarten to a master craftsman.
Stage 1: The "Playground" Phase (Learning to Move)
Before learning specific jobs, the robot is given unstructured "play" data.
- The Analogy: Imagine a child playing with blocks, dropping them, picking them up, and stacking them randomly without a teacher telling them what to do.
- Why it helps: The robot learns the physics of moving things. It learns how to recover if it drops a block or how to grab something awkwardly. This gives it a "gut feeling" for how objects behave, so it doesn't panic when things go slightly wrong.
Stage 2: The "Repetition" Phase (Learning to Endure)
Most robots are trained on "Success then Stop" videos. Habilis-β is trained on cyclic data.
- The Analogy: Instead of watching a video of someone making one sandwich and stopping, Habilis-β watches a video of someone making sandwiches for 10 hours straight. It sees the bread get stale, the table get messy, and the person getting tired.
- Why it helps: It learns to handle drift. In the real world, things aren't perfectly placed every time. This training teaches the robot to keep going even when the starting position is slightly off, preventing it from needing a human reset.
Stage 3: The "Speed Shaping" Phase (ESPADA)
This is the secret sauce for speed. The robot uses a technique called ESPADA (Spatially Aware Downsampling).
- The Analogy: Imagine a video of a person walking across a room.
- Walking across the empty room: The robot learns to fast-forward this part. It doesn't need to see every step; it just needs to know "I'm going there."
- Picking up a fragile egg: The robot slows down to slow-motion. It needs to see every tiny movement to be precise.
- Why it helps: It stops the robot from being slow and boring when it's safe to be fast. It compresses the "boring" parts of the movement so the robot can zip across the room and only focus on the tricky parts.
3. The Hardware: The "On-Device" Brain
Many robots rely on the cloud (the internet) to think. If the internet lags, the robot lags.
- The Analogy: Habilis-β is like a smartphone with a powerful processor that does all the math right in your pocket. It doesn't need to ask Siri or Google for help.
- Why it helps: Because it thinks locally, it can react instantly. It can make decisions in milliseconds, allowing it to correct mistakes before they become disasters.
4. The "Volume Knob" (CFG)
Finally, the system has a deployment-time knob called Classifier-Free Guidance (CFG).
- The Analogy: Think of this as a volume knob for instructions.
- Low Volume: The robot relies more on its own "gut feeling" (the play data) to keep moving smoothly.
- High Volume: The robot follows the human's specific instructions very strictly, even if it means moving more aggressively.
- Why it helps: You can tune the robot depending on the job. If you need it to be super precise, you turn the knob up. If you need it to be fast and flexible, you turn it down.
The Results: The "Factory Veteran"
When tested against other top robots (like and GR00T):
- In Simulation: Habilis-β did 572 tasks per hour, while the others did about 120. It also went much longer without needing a human to fix it.
- In the Real World: On a real humanoid robot doing a logistics job, Habilis-β completed 124 tasks per hour (compared to 19 for the others) and ran for 137 seconds between needing help (compared to 46 seconds for the others).
Summary
Habilis-β is a robot brain that stops acting like a nervous student taking a single test and starts acting like a seasoned worker. It learns by playing, it learns by working long shifts, it speeds up when it's safe, and it thinks fast enough to catch its own mistakes. It's designed not just to be "smart," but to be productive and reliable in the real world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.