← Latest papers
💻 computer science

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use

This paper introduces **In-Context VLA**, a framework that enhances Vision-Language-Action models by replacing free-form chain-of-thought generation with in-context post-training on structured perceptual evidence and agentic tool use, thereby enabling grounded language consumption for superior low-level control and state-of-the-art performance across simulation and real-world manipulation tasks.

Original authors: Jiarui Yang, Wen Huang, Jiale Zhang, Maowei Hu, Hang Guo

Published 2026-08-07
📖 3 min read☕ Coffee break read

Original authors: Jiarui Yang, Wen Huang, Jiale Zhang, Maowei Hu, Hang Guo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to do chores, like making a sandwich or tidying a messy room. For a long time, scientists have tried to teach these machines by showing them videos of humans doing the job and telling the robot, "Copy exactly what you see." This works okay for simple tasks, but it's like trying to learn a new language by only memorizing a single, rigid sentence. If you ask the robot to "pick up the red mug" but then say, "grab that crimson cup," the robot might get confused because it never learned to understand the meaning behind the words, only to match the specific sounds it heard during training.

To fix this, researchers recently tried giving robots a "thinking" step, similar to how humans talk to themselves before acting. They hoped the robot would first write out a plan like, "Okay, the cup is on the table, and I need to move my hand to the left," and then act. This is called "Chain-of-Thought." But here's the twist: for robots, talking too much before moving actually makes them worse at their jobs. It's like trying to drive a car while writing a novel about the road; by the time you finish your sentence, you've missed the turn. The robot gets stuck in a loop of narrating its own thoughts instead of actually grabbing the object, and it often makes up facts about where things are because it's guessing rather than looking.

This paper introduces a new way to teach robots called VLA-Talker. Instead of forcing the robot to write its own thoughts, the researchers give the robot a "smart assistant" that does the looking and reporting for it. Think of it like a robot with a super-powered pair of glasses and a helpful co-pilot. When the robot needs to know where a spoon is, it doesn't guess or write a paragraph about it. Instead, it instantly asks its co-pilot (which uses special tools like depth cameras and object detectors) to say, "The spoon is 42 pixels to the right and slightly lower than your hand." The robot then reads this clear, factual report and acts immediately.

The researchers found that this method is a game-changer. By letting the robot consume clear, grounded language from its tools instead of generating its own confusing chatter, the robot becomes much faster and more accurate. In their tests, this approach didn't just work better; it was nearly 4.6 times faster than the old "thinking" methods because the robot didn't waste time writing essays before moving. They tested this in complex computer simulations and even on a real human-shaped robot, and it consistently solved tasks that other robots struggled with, proving that sometimes, the best way to think is to listen to someone else who knows the facts.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →