← Latest papers
💻 computer science

Learning Category-level Last-meter Navigation from RGB Demonstrations of a Single-instance

This paper presents an object-centric imitation learning framework that enables a quadruped mobile manipulator to achieve precise, category-level last-meter navigation for manipulation-ready positioning using only RGB observations and text prompts, without relying on depth sensors, LiDAR, or map priors.

Original authors: Tzu-Hsien Lee, Fidan Mahmudova, Karthik Desingh

Published 2026-07-02
📖 5 min read🧠 Deep dive

Original authors: Tzu-Hsien Lee, Fidan Mahmudova, Karthik Desingh

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to pick up a specific chair in a messy living room. You might think, "Just tell the robot to go to the chair!" But here's the problem: most robots are like people with bad depth perception. They can get you to the neighborhood of the chair (within about 3 feet), but they can't get you to stand in the exact right spot to grab it. If the robot is even a little off-center or facing the wrong way, the robot's arm (manipulator) will miss the chair entirely, just like you trying to grab a cup while your hand is slightly too far to the left.

This paper introduces a solution for that final, tricky step: the "last-meter navigation." It's the difference between parking your car in the right neighborhood versus pulling it perfectly into the garage spot so you can walk straight out the door.

Here is how they did it, explained simply:

1. The Problem: "Good Enough" Isn't Good Enough

Most robot navigation systems are trained to stop when they are roughly 1 meter away from a target. That's fine for walking up to a door, but terrible for picking things up. The authors call this the "gap." If the robot doesn't get the positioning perfect, the rest of the task fails.

Usually, to get that perfect precision, robots need expensive tools like 3D lasers (LiDAR) or detailed maps of the room. The authors wanted to know: Can a robot do this using only a standard camera (RGB), just like a human uses their eyes?

2. The Solution: Learning by Watching (Imitation)

Instead of programming the robot with complex rules (like "if the chair looks this big, turn left"), they taught the robot to learn by watching.

  • The Teacher: They used a real robot (a Boston Dynamics Spot) to record a human driving it to a specific green chair. The robot recorded what the camera saw and exactly how it moved to get there.
  • The Student: They trained a computer model to mimic these movements.
  • The Secret Sauce: They didn't just show the robot the whole room. They taught it to focus specifically on the object (the chair). They used a "text prompt" (like saying "find the chair") to tell the robot which object to look at, then used a special filter to ignore everything else in the room.

3. How the Robot "Thinks"

The robot's brain works in three steps, similar to how you might navigate:

  1. Spot the Target: The robot looks at its current view and the "goal" view (a picture of what the chair should look like when it's in the perfect spot). It uses a text command to find the chair in both pictures.
  2. Compare the Views: It creates a "score matrix." Imagine holding two transparent sheets with a grid on them. One sheet is the current view, the other is the goal view. The robot slides them over each other to see how the patterns match up. If the chair in the current view is to the left of where it should be, the robot knows to move right.
  3. Move and Stop: The robot takes small steps (forward, sideways, or turning) to align the patterns. Crucially, they added a "brake" system. Since the robot's training data had some tiny wobbles at the end, the robot might not know exactly when to stop. So, they added a rule: "If the chair looks 60% similar to the goal picture and is centered, hit the brakes."

4. The Results: One Chair, Many Chairs

The most impressive part is how well this generalizes.

  • The Training: They only trained the robot on one single green chair in one specific room.
  • The Test: They then tested it on completely different chairs (brown, wooden, metal) in different rooms (living rooms, offices, and even outside).
  • The Outcome: The robot successfully navigated to the correct spot about 75% to 90% of the time, depending on how strict the test was. It worked even without 3D lasers or maps, using only the camera feed.

5. Where It Struggles

The paper is honest about the limitations. The system is very sensitive to lighting.

  • In bright, natural light (like outdoors), the robot performed best because the camera could see the chair clearly.
  • In dimly lit rooms, the robot sometimes got confused because it couldn't clearly "see" the chair to align itself.
  • If the chair is blocked by something else (occlusion), the robot can't do its job because it relies on seeing the target clearly.

Summary

This paper proves that a robot can learn to park itself perfectly next to an object it has never seen before, using only a standard camera and a little bit of "show-and-tell" training. It bridges the gap between "getting close" and "being ready to grab," all without needing expensive 3D sensors. It's like teaching a robot to drive into a parking spot by showing it one example, and then letting it figure out how to park in any other spot on its own.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →