← Latest papers
💻 computer science

SkillWrapper: Generative Predicate Invention for Task-level Robot Planning

This paper introduces SkillWrapper, a method that leverages foundation models to automatically invent formal, provably sound and complete symbolic predicates from raw RGB observations, enabling robots to generalize black-box skills into abstract representations for solving unseen long-horizon tasks.

Original authors: Ziyi Yang, Benned Hedegaard, Ahmed Jaafar, Yichen Wei, Skye Thompson, Shreyas S. Raman, Haotian Fu, Stefanie Tellex, George Konidaris, David Paulius, Naman Shah

Published 2026-06-09
📖 5 min read🧠 Deep dive

Original authors: Ziyi Yang, Benned Hedegaard, Ahmed Jaafar, Yichen Wei, Skye Thompson, Shreyas S. Raman, Haotian Fu, Stefanie Tellex, George Konidaris, David Paulius, Naman Shah

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a robot that is incredibly strong and dexterous at the physical level. It can pick up a teapot, pour tea, and stack a bowl on a plate. However, right now, this robot is like a brilliant pianist who only knows how to press individual keys but doesn't understand music theory. It can play a single note, but it doesn't know how to compose a song or plan a whole concert.

This is the problem SkillWrapper solves.

The Problem: The "Black Box" Robot

Currently, many robots come with "black box" skills. These are pre-programmed actions (like "Pick up object" or "Pour liquid") that work well on their own. But if you want the robot to do a complex, multi-step task—like "Make a cup of tea"—it doesn't know how to string these skills together.

To get a robot to plan a complex task, humans usually have to manually write a rulebook for the robot. They have to explain: "To pour tea, the robot must first be holding the teapot, and the cup must be empty." This is tedious, slow, and impossible to do for every new robot or environment.

The Solution: Teaching the Robot to Write Its Own Rulebook

The researchers created a system called SkillWrapper. Instead of a human writing the rulebook, SkillWrapper teaches the robot to invent its own "vocabulary" and "rules" by watching itself try things out.

Here is how it works, using a simple analogy:

1. The "Trial and Error" Phase (Active Data Collection)

Imagine a robot in a kitchen. It doesn't know the rules yet. SkillWrapper tells the robot: "Go try picking up the teapot. Now try picking up the sponge. Now try stacking the bowl."

  • Success: The robot picks up the teapot.
  • Failure: The robot tries to pick up the sponge, but it's too slippery, and it drops it.

The robot records these moments: "I tried to pick up the teapot, and it worked. I tried to pick up the sponge, and it failed."

2. The "Aha!" Moment (Generative Predicate Invention)

This is the magic part. The robot looks at its successes and failures and asks: "Why did one work and the other fail?"

In the past, computers needed humans to tell them the answer (e.g., "The sponge is too heavy"). But SkillWrapper uses a Foundation Model (a super-smart AI that understands images and language, like a very advanced version of ChatGPT or DALL-E).

  • The robot shows the AI two pictures: one of a successful pickup and one of a failed pickup.
  • The AI looks at the images and invents a new word (a predicate) to describe the difference.
  • The AI says: "Ah! The difference is that the teapot has a handle, but the sponge is slippery. Let's invent a new rule called GripperEmpty or ObjectIsStable."

The robot creates a new, human-readable concept that explains why the action worked or failed. It's like the robot suddenly learning the word "slippery" and realizing, "Oh, I can't pick up slippery things with my current grip!"

3. Building the Rulebook (Operator Learning)

Once the robot has invented these new words (predicates), it starts building a logical map.

  • Rule: "To Pour, you must be Holding a Teapot and the Cup must be Empty."
  • Rule: "To Stack, the Plate must be Flat."

The robot groups these rules together to form a Symbolic Model. This is a high-level map of the world that ignores the messy details (like the exact angle of the arm) and focuses on the logic (like "is the cup empty?").

4. The Grand Plan (Task-Level Planning)

Now, when you give the robot a complex goal like "Make tea," it doesn't panic. It uses a standard planning algorithm (the same kind of logic used in video game AI or logistics software) to look at its new rulebook.

  • It sees the goal: "Cup has tea."
  • It checks the rules: "I need to Pour."
  • It checks the preconditions: "Do I have a teapot? Is the cup empty?"
  • It builds a step-by-step plan: Pick up Teapot -> Move to Cup -> Pour.

Why This Paper is Special

The paper claims three major things that make this different from previous attempts:

  1. It starts from zero: Unlike other methods that need humans to give them a starting list of rules, SkillWrapper starts with nothing. It invents the vocabulary from scratch just by watching the robot move.
  2. It's mathematically guaranteed to be correct: The authors didn't just guess that this would work. They proved with math that if the robot follows this process, the rules it learns will be Sound (it won't plan impossible things) and Complete (it won't miss a solution if one exists). It's like proving a bridge is safe to cross before letting anyone walk on it.
  3. It works in the real world: They tested this on real robots (not just simulations). The robots learned to plan complex tasks, like stacking objects or pouring liquids, using only camera images as input.

The Bottom Line

SkillWrapper is a system that lets a robot teach itself the "grammar" of its own actions. Instead of a human writing a manual, the robot tries things, fails, asks a smart AI "Why did that fail?", invents a new word to describe the problem, and then uses that word to plan its next move. This allows robots to solve long, complicated tasks in the real world without needing a human expert to program every single rule.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →