← Latest papers
💻 computer science

HABIT: Human-Aware Behavior and Interaction Training Dataset for Robot Manipulation

The paper introduces HABIT, a large-scale dataset of over 10,000 human-present robot manipulation episodes organized into Collaborator, Coworker, and Supervisor roles, which enables policies to learn essential human-aware behaviors like synchronization, yielding, and gesture grounding that are absent in traditional human-absent datasets.

Original authors: Jaehwi Song, Suchae Jeong, Byeongguk Jeon, Sungdong Kim, Minjoon Seo, Hyungmok Son, Kimin Lee

Published 2026-07-01
📖 5 min read🧠 Deep dive

Original authors: Jaehwi Song, Suchae Jeong, Byeongguk Jeon, Sungdong Kim, Minjoon Seo, Hyungmok Son, Kimin Lee

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot how to do chores. Most robots today are trained in a "ghost town" scenario: they practice moving boxes, wiping tables, and sorting trash in an empty room where no humans are around. They become very good at the physical motions, but they have no idea how to act when a real person is standing right next to them. They might bump into you, ignore your hand gestures, or try to grab a tool while you are still holding it.

The paper introduces HABIT (Human-Aware Behavior and Interaction Training), a new dataset designed to fix this. Think of HABIT not just as a library of robot moves, but as a library of dance steps where the robot and a human partner are learning to move together.

The Three "Dance Styles"

The researchers organized the tasks into three distinct roles, like different genres of dance, to teach the robot how to handle different social situations:

  1. The Collaborator (The Tango):

    • The Scenario: The human and robot must work together on the same object at the same time.
    • The Analogy: Imagine two people trying to carry a heavy couch up a staircase. If one person moves too fast or too slow, they trip. In this role, the robot learns to wait for the human to lift a corner before it moves, or to hand over a tool exactly when the human reaches for it. It's all about synchronization.
    • Example: A human places a napkin on a tray, and the robot lifts the tray to help.
  2. The Coworker (The Conga Line):

    • The Scenario: The human and robot are in the same room doing different tasks, but they share the same space.
    • The Analogy: Imagine two people cooking in a small kitchen. One is chopping vegetables, and the other is stirring a pot. They aren't touching each other, but they have to dodge around each other so they don't collide. The robot learns to yield (step aside) if the human walks into its path, rather than plowing forward.
    • Example: A human sorts trash on one side of a table while the robot packs boxes on the other side, constantly moving out of each other's way.
  3. The Supervisor (The Conductor):

    • The Scenario: The human gives directions, and the robot follows.
    • The Analogy: Imagine a conductor waving a baton to tell an orchestra when to play. The robot learns to watch the human's hands and eyes. If the human points at a specific donut, the robot doesn't just grab the first one it sees; it looks at the finger and grabs that specific donut. It learns gesture grounding.
    • Example: A human points to a specific container, and the robot puts bread inside it.

How They Collected the Data

To make sure the robot actually learned these social skills, the researchers didn't just let the human and robot play around randomly. They used a strict "script" (called a workflow) that forced the human and robot to react to each other in real-time.

  • No "Cheat Codes": The human operator couldn't just guess what the robot would do next. They had to wait for the robot to move before reacting, and vice versa. This ensured the robot learned to respond to actual human cues, not just memorized patterns.
  • Safety First: If the robot looked like it was about to bump into the human, the human operator would immediately pull the robot back. This taught the robot that "stopping" is a valid and important move.

What Happened When They Tested It?

The researchers took two existing robot brains (AI models) and trained them on this new HABIT data. They compared these "Habit-trained" robots against robots trained only on the old "ghost town" data.

  • The Results: The HABIT-trained robots were significantly better.
    • In the Coworker role, they stopped crashing into humans.
    • In the Supervisor role, they actually looked at where the human pointed instead of guessing.
    • In the Collaborator role, they timed their moves perfectly with the human.
  • The "Superpower": The most interesting finding was sample efficiency. When they tried to teach the HABIT-trained robots a new task they hadn't seen before, they learned it much faster (using fewer examples) than the robots trained on the old data. It's as if HABIT taught the robot the concept of "being polite and aware," which made learning new specific chores much easier.

The Bottom Line

The paper argues that for robots to be useful in our homes and offices, they can't just be strong and precise; they have to be socially aware. By training robots on data where humans are actually present and interacting, we get robots that don't just work around us, but work with us.

Limitations: The authors admit this was done in a controlled lab with a specific type of robot and a limited number of human operators. While they varied the humans' clothes and body types to test adaptability, real-world environments are messier than their setup. However, the core message stands: robots need to see humans to learn how to be good partners.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →