← Latest papers
💻 computer science

LocalNav: Distilling Frontier VLMs and Embodied RL for On-Device Object Goal Navigation

This paper presents "LocalNav," a framework that distills complex spatial-semantic reasoning from frontier Vision Language Models into a lightweight 4B-parameter model and optimizes it with E-RLVR and quantization, enabling high-performance, low-latency object goal navigation on resource-constrained edge devices without relying on cloud execution.

Original authors: Nicolas Baumann, Liam Boyle, Pu Deng, Edoardo Ghignone, Boyang Sun, Marc Pollefeys, Luca Benini, Michele Magno

Published 2026-06-29
📖 4 min read☕ Coffee break read

Original authors: Nicolas Baumann, Liam Boyle, Pu Deng, Edoardo Ghignone, Boyang Sun, Marc Pollefeys, Luca Benini, Michele Magno

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, but very heavy, robot brain. This brain is like a giant library of knowledge that can understand complex instructions like, "Find the drawers under the microwave." The problem is, this brain is so big and heavy that it can't fit inside a small robot; it has to live in a giant cloud computer far away. Every time the robot needs to make a decision, it has to send a message to the cloud, wait for a reply, and then move. This is slow, expensive, and doesn't work if the robot is in a place with no internet.

The paper introduces LocalNav, a new way to shrink that giant brain down so it can fit inside a small robot (like one with a Jetson Orin chip) and work entirely on its own, without needing the cloud.

Here is how they did it, using three simple steps:

1. The "Shadow Student" (Distillation)

Think of the giant cloud brain (like Claude or GPT) as a Master Chef. The Master Chef can cook a perfect, complex meal, but it takes a long time and uses a huge kitchen. The researchers wanted to teach a Student Chef (a much smaller, 4-billion-parameter model called Qwen3.5-4B) how to cook the same dishes.

Instead of forcing the student to read millions of cookbooks, they just showed the student about 500 examples of the Master Chef solving navigation puzzles. They said, "Watch how the Master Chef thinks: 'I see a microwave, so the drawers must be below it.'" By copying these specific examples, the Student Chef learned to navigate just as well as the Master Chef, but with a brain small enough to fit in a backpack.

  • The Result: The small robot brain achieved a success rate of 34.5%, which is very close to the giant cloud brain's 39.7%.

2. The "Brevity Coach" (E-RLVR)

There was one catch: The Student Chef was a bit too chatty. When asked to find the drawers, the Master Chef might say, "I see a silver appliance that looks like a microwave, and I recall that drawers are usually below microwaves, so I will go there." That's a lot of words to type! For a robot with limited battery and processing power, typing all those words takes too long.

To fix this, the researchers used a technique called E-RLVR (Embodied Reinforcement Learning from Verifiable Rewards). Imagine a strict coach who watches the student robot.

  • If the robot takes a long time to answer or says too many words, the coach gives a "bad grade."
  • If the robot finds the object quickly and says, "Drawers found. Done!" in just a few words, the coach gives a "gold star."

The robot learned to "do" rather than "talk." It figured out how to get the job done using the fewest words possible. This cut the time it took to think and speak by 71.8%.

3. The "Compact Suit" (Quantization)

Finally, to make the robot even faster, they put the brain in a "compact suit." This is a technical trick called quantization, which is like compressing a high-resolution photo into a smaller file size without losing the important details. This made the robot's brain run even faster on its small hardware.

The Final Outcome

By combining these three tricks:

  1. Learning from a Master (Distillation),
  2. Learning to be concise (E-RLVR), and
  3. Compressing the brain (Quantization),

The researchers created a robot that can understand complex instructions like "Find the person sitting at the table" and navigate a real apartment all by itself, without needing the internet.

In their real-world test, they sent a handheld robot into an apartment with four different missions: find a sofa, a piano, a bathtub, and a person. The robot successfully built a mental map of the rooms, found the objects, and stopped exactly where it needed to, proving that a small, local brain can do the work of a giant cloud brain.

In short: They took a super-smart but slow cloud brain, taught a small local brain how to think like it, taught the small brain to stop talking so much, and squeezed it into a tiny package. Now, the robot can think and move fast, right on the device.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →