Ask When It Pays: Cost-Aware Open-Ended Interaction for Instance Goal Navigation
This paper addresses the inefficiency of prior interactive navigation methods by reframing Instance Goal Navigation as a cost-sensitive uncertainty-reduction problem, introducing a new benchmark with a cost-aware metric and a zero-shot MLLM navigator that selectively queries an oracle only when the expected information gain justifies the interaction cost.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a robot tasked with finding a specific object in a house, like "the blue mug." But here's the catch: there are ten blue mugs on different tables, and you don't know which one your boss actually wants. This is the problem of Instance Goal Navigation.
The paper argues that while robots can ask for help, they often ask the wrong questions or ask too many of them, wasting time and energy. The authors propose a new way for robots to interact: "Ask When It Pays."
Here is a breakdown of their solution using simple analogies:
1. The Problem: The "Free Advice" Trap
Imagine you are lost in a massive mall. You can ask a helpful stranger (the "Oracle") for directions.
- Old Way: Previous robot methods treated all questions as free. So, a robot might ask, "Is it the red one?" "Is it the blue one?" "Is it on the left?" "Is it on the right?" It might ask 20 questions to find the answer, boosting its success rate but taking forever.
- The Issue: Asking a question that gives you a huge clue (like "It's in the kitchen") is very valuable. Asking a question that gives a tiny clue (like "Is it shiny?") is less valuable. If the robot doesn't know the difference, it wastes its "interaction budget" on cheap questions.
2. The Solution: The "Cost-Aware" Strategy
The authors treat asking a question like spending money.
- The Price Tag: They analyzed thousands of human navigation logs to figure out which types of questions are most helpful. They assigned a "cost" to each type:
- Appearance Questions ("Is it red?"): Low cost.
- Location Questions ("Is it in the kitchen?"): Medium cost.
- Route Questions ("Go left at the hall"): High cost (because this is a big hint).
- Confirmation Questions ("Is this the one?"): Low cost.
- The Strategy: The robot is now programmed to only ask a question if the value of the answer is greater than the cost of asking. It's like deciding whether to buy a lottery ticket: only buy it if the potential prize is worth the price of the ticket.
3. The New Robot: TANDEM
The authors built a robot named TANDEM to test this idea. Think of TANDEM as a team of two people working together:
- The Planner (The Brain): This is a large AI that looks at the room, remembers where it's been, and decides what to do. It decides whether to move, stop, or ask a question. It doesn't know exactly how to walk; it just knows the goal.
- The Grounder (The Legs): This part takes the Planner's vague idea (e.g., "Go to that chair") and turns it into actual physical steps, making sure the robot doesn't bump into walls.
How TANDEM asks:
Instead of asking random questions, TANDEM asks specific types of questions at the right time:
- Early on: It asks about Appearance or Region to narrow down which mug it is looking for.
- In the middle: If it gets confused at a hallway intersection, it asks a Route question to get a general direction (e.g., "The room is past the dining table").
- At the end: It asks Confirmation questions to make sure it hasn't stopped at the wrong mug.
4. The Results: Smarter, Not Harder
The researchers created a new test environment (a digital simulation) where they could control how many "fake" mugs (distractors) were in the room to make the task harder.
- The Finding: Robots that asked questions freely (the "Naive" approach) often got confused or wasted time. TANDEM, which only asked when it "paid off," was much more efficient.
- The "Aha" Moment: The paper found that TANDEM asks the most questions when the robot is truly confused (like at a complex hallway junction). It doesn't ask just to be chatty; it asks to solve a specific puzzle.
- The Metric: They introduced a new score called Weighted Success Rate. This score doesn't just count if the robot found the object; it subtracts points for every expensive question asked. TANDEM won because it found the object and spent the least amount of "interaction money."
Summary
The paper teaches robots a valuable lesson: Don't just ask for help; ask for the right help at the right time. By treating questions as a limited resource with different prices, the robot becomes a smarter navigator that solves problems efficiently rather than just guessing its way through.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.