← Latest papers
💻 computer science

Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models

The paper proposes Latent Reasoning VLA (LaRA-VLA), a unified framework that internalizes multi-modal chain-of-thought reasoning into continuous latent representations to achieve efficient, real-time embodied control with up to 90% reduced inference latency compared to explicit reasoning approaches.

Original authors: Shuanghao Bai, Jing Lyu, Wanqi Zhou, Zhe Li, Dakai Wang, Lei Xing, Xiaoguang Zhao, Pengwei Wang, Zhongyuan Wang, Cheng Chi, Badong Chen, Shanghang Zhang

Published 2026-05-12
📖 4 min read☕ Coffee break read

Original authors: Shuanghao Bai, Jing Lyu, Wanqi Zhou, Zhe Li, Dakai Wang, Lei Xing, Xiaoguang Zhao, Pengwei Wang, Zhongyuan Wang, Cheng Chi, Badong Chen, Shanghang Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to make a sandwich.

The Old Way (Explicit Chain-of-Thought):
Currently, most advanced robot brains work like a student who is forced to write out every single thought in a diary before taking a single action.

  • The Robot sees: "I need a sandwich."
  • The Robot thinks (out loud): "Okay, first I need to find the bread. I see a loaf on the left. Now I need to grab the knife. I see the knife on the right. I need to move my arm slowly..."
  • The Robot acts: Only after writing this entire paragraph does it move its arm.

The Problem: This is incredibly slow. By the time the robot finishes writing its "diary entry," the sandwich is already cold, or the robot has missed its chance to grab the bread. Also, writing in words (discrete tokens) is a bit clunky for a robot that needs to move its arm in smooth, continuous curves. It's like trying to drive a car by only allowing yourself to move in 1-inch jumps.

The New Way (LaRA-VLA):
The paper introduces LaRA-VLA (Latent Reasoning VLA). Think of this as teaching the robot to think in its head without writing anything down.

Instead of spitting out a long paragraph of text, the robot internalizes its reasoning into a compact, continuous "feeling" or "vibe" (a latent representation).

  • The Robot sees: "I need a sandwich."
  • The Robot thinks (silently): It instantly processes the location of the bread, the knife, and the path its arm needs to take, all in a single, smooth mental snapshot.
  • The Robot acts: It moves immediately.

The Three-Step Training Camp (Curriculum Learning)

The paper explains that you can't just tell a robot to "think silently" right away. It needs a training camp with three stages:

  1. Stage 1: The "Show Your Work" Phase.
    The robot is forced to write out its thoughts (text) and predict what the next picture will look like (visuals). It learns the rules of the game by explicitly stating, "I am moving the cup to the left."

  2. Stage 2: The "Whisper" Phase.
    The trainers start replacing the robot's written sentences with "whispers" (latent tokens). At first, the robot writes half the sentence and whispers the rest. Slowly, it writes less and whispers more. It learns to compress its complex thoughts into a tiny, efficient mental package.

  3. Stage 3: The "Silent Master" Phase.
    The robot no longer writes or whispers anything out loud. It has fully internalized the reasoning. It looks at the scene, processes the "whisper" in its mind, and immediately executes the action.

Why is this better?

  • Speed: Because the robot isn't wasting time typing out a diary entry, it can react 90% faster. The paper claims this drops the thinking time from over 7 seconds down to just 0.13 seconds (135 milliseconds). That's fast enough for real-time control.
  • Smoothness: Since the robot thinks in "continuous feelings" rather than "discrete words," its movements are smoother and match how the physical world actually works.
  • Robustness: The paper tested the robot with blurry or noisy camera feeds (like looking through a foggy window). The "silent thinkers" (LaRA-VLA) handled the mess much better than the "diary writers," suggesting their internal mental models are more stable.

The Result

In tests, this new method beat almost all other top-tier robot models. It was better at long, complicated tasks (like sorting fruit or stacking bowls) and did so much faster.

In short: LaRA-VLA teaches robots to stop "talking" their way through a task and start "feeling" their way through it, resulting in a robot that is both a genius planner and a lightning-fast worker.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →