← Latest papers
💻 computer science

Hierarchical Prompting with Dual LLM Modules for Robotic Task and Motion Planning

This paper presents a hierarchical language-driven framework that integrates a high-level LLM planning agent with a low-level spatial reasoning sub-module and object detection tools to achieve an 86% success rate in robotic task and motion planning across diverse service scenarios.

Original authors: Karolina Źróbek, Tessa Pulli, Paweł Gajewski, Antonio Galiza Cerdeira Gonzalez, Bipin Indurkhya

Published 2026-06-30
📖 4 min read☕ Coffee break read

Original authors: Karolina Źróbek, Tessa Pulli, Paweł Gajewski, Antonio Galiza Cerdeira Gonzalez, Bipin Indurkhya

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to help you in your kitchen. If you just say, "Make me a fruit bowl," a standard robot might get confused. It knows what a "fruit" is, but it doesn't know where the fruit is, how to grab it without squishing it, or exactly where to put it next to the bowl.

This paper presents a new way to talk to robots using a "two-brain" system. Instead of asking one super-smart computer to do everything at once, the researchers split the job between two specialized Large Language Models (LLMs), which are like very advanced AI chatbots.

Here is how their system works, using simple analogies:

The Two-Brain System

1. The "General Manager" (High-Level Agent)
Think of this AI as the project manager. You give it a natural command like, "Clear the table of all fruits."

  • What it does: It breaks the big goal down into a to-do list. It decides, "First, I need to find the apples. Then, I need to pick them up. Then, I need to put them in the trash."
  • How it thinks: It uses a method called "ReAct" (Reason + Act). It thinks, "I need an apple," then it asks the robot's eyes to look for one, sees the apple, and then says, "Okay, grab it." It keeps a running log of what it has done so it doesn't forget.
  • The limitation: While this manager is great at logic, it is terrible at math. If you tell it to "put the cup to the left of the plate," it might understand the words, but it doesn't actually know the exact 3D coordinates to move the robot arm to. It might guess, and that guess could be physically impossible.

2. The "Spatial Architect" (Low-Level Sub-Module)
This is the specialized engineer who handles the geometry. The General Manager doesn't talk to the robot arm directly; it talks to this Architect.

  • What it does: When the Manager says, "Place the mug next to the plate," the Architect takes over. It looks at the 3D shapes of the mug and the plate (using a camera system called YOLOX-GDRNet) and calculates the exact math: "The plate is at coordinate X, Y, Z. To be 'next to' it, the mug needs to be at X + 5cm."
  • Why separate them? If you ask the General Manager to do the math and the planning at the same time, it gets overwhelmed and makes mistakes (like trying to put a mug inside a solid table). By separating the "what to do" from "exactly where to put it," the system is much more reliable.

How They Tested It

The researchers put this system through 24 different tests in a simulation and on a real robot (a PAL Robotics Tiago arm).

  • The Results: The system succeeded in 86% of the tasks.
  • Simple vs. Hard: It was very good at simple commands (like "pick up the banana") and even better at recognizing impossible tasks (like "put the banana inside the solid table"). In those impossible cases, the robot correctly said, "I can't do that," 100% of the time.
  • Human Feedback: When humans looked at the final results, they gave it a score of about 6.9 out of 10. They liked it when the robot did clear tasks (like stacking dishes), but they were less sure when the instructions were vague (like "make a salty snack"), because "salty snack" is open to interpretation.

The Real-World Test

They also tried this on a real robot arm in a real room.

  • Success: The robot was great at finding objects (like a mustard bottle or an apple) even when the table was messy. It figured out the plan perfectly.
  • Failures: The few times it failed, it wasn't because the AI got confused about the plan. It failed because of physical hardware issues, like the camera being slightly misaligned or the robot arm bumping into something. The "brain" was working; the "body" just had a minor glitch.

The Bottom Line

The paper shows that by splitting the job into a Manager (who handles the logic and conversation) and an Architect (who handles the 3D math and placement), robots can understand human instructions much better.

However, the authors are honest about the limits: while the robot is getting smarter at understanding language and space, it still struggles with the messy reality of the physical world. It's a big step forward for making robots that can actually help us at home without needing us to speak "robot code."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →