← Latest papers
💻 computer science

OneVLA: A Unified Framework for Embodied Tasks

OneVLA introduces a unified framework that integrates navigation and manipulation into a single architecture with a shared action head and a multi-stage progressive training strategy, achieving state-of-the-art performance in both simulated and real-world environments to advance the development of general-purpose robotic agents.

Original authors: Lingfeng Zhang, Xiaoshuai Hao, Yingbo Tang, Lei Zhou, Shuyi Zhang, Jinkun Liu, Hongsheng Li, Chenhao Zhang, Qiang Zhang, Hangjun Ye, Xiaojun Liang, Long Chen, Wenbo Ding

Published 2026-06-02
📖 4 min read☕ Coffee break read

Original authors: Lingfeng Zhang, Xiaoshuai Hao, Yingbo Tang, Lei Zhou, Shuyi Zhang, Jinkun Liu, Hongsheng Li, Chenhao Zhang, Qiang Zhang, Hangjun Ye, Xiaojun Liang, Long Chen, Wenbo Ding

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a robot assistant. Until now, building this robot was like hiring two separate specialists: one person who is an expert at walking (navigation) but terrible at using their hands, and another person who is a master chef (manipulation) but can't walk across the room to get ingredients. If you wanted the robot to "walk to the kitchen and put a cup in the microwave," you had to switch between these two different "brains" or train two separate models.

The paper introduces OneVLA, which is like hiring a single "Swiss Army Knife" employee who is naturally good at both walking and using their hands, all within one single brain.

Here is a simple breakdown of how they did it:

1. The "Universal Remote Control" (Unified Action Head)

Most robots have different "remotes" for different jobs. One remote has buttons for moving forward or turning left. Another remote has sliders for moving a robotic arm up, down, or twisting.

OneVLA creates a single, universal remote.

  • It combines the buttons for walking and the sliders for the arm into one long strip of controls.
  • When the robot gets an instruction like "Go to the door," it only presses the "walking buttons" on the strip.
  • When the instruction is "Pick up the cup," it only moves the "arm sliders."
  • The Magic: The robot doesn't need to swap out its brain or its remote. It just knows which part of the remote to use based on the task.

2. The "Three-Step Training Camp" (Multi-Stage Strategy)

You can't just throw a robot into a complex world and expect it to know everything immediately. The authors trained OneVLA in three specific phases, like a student progressing through school:

  • Stage 1: Learning to Use Hands. First, they taught the robot how to grab, lift, and place objects using a robotic arm. They also taught it to look at pictures and describe them (so it understands the world).
  • Stage 2: Learning to Walk. Next, they added navigation data. Now the robot learns how to move from point A to point B. Because it already knows how to "see" and "think" from Stage 1, it learns to walk faster and smarter.
  • Stage 3: Learning to Think Aloud (Chain-of-Thought). Finally, they taught the robot to "talk through" its thinking process. Before it moves, it says to itself, "I am in the hallway. I need to turn right to get to the kitchen." This step helps the robot connect the dots between what it sees, what it needs to do, and how to move.

3. The Result: Two Skills, One Brain

The paper claims that by training these two skills together, they actually help each other get better.

  • The Analogy: Think of it like a basketball player who also plays soccer. Learning to dribble a basketball (manipulation) might improve their footwork and balance, which helps them run better on the soccer field (navigation).
  • The Proof: In tests, this single robot (OneVLA) beat robots that were specialized only in walking or only in using hands. It was also better than other "hybrid" robots that tried to do both but still required separate "modes" or settings for each task.

In Summary

OneVLA is a new type of robot brain that doesn't need to be reprogrammed when you ask it to walk or when you ask it to grab something. It uses one unified system to understand language, see the world, and control both its feet and its hands. The authors show that teaching these skills together makes the robot smarter at both than if you taught them separately.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →