← Latest papers
💻 computer science

Teaching is a Process: The TOSS Framework for Modeling Human Teaching Decisions in Human-Interactive Robot Learning

This paper introduces the TOSS Framework, a theoretical model derived from an exploratory study of human teaching behaviors, which conceptualizes human-robot teaching as a procedural loop of triggers, signals, objectives, and strategies to better align robot learning with human intent and facilitate the design of more effective, human-centered teaching systems.

Original authors: Bernhard Hilpert, Kim Baraka, Joost Broekens

Published 2026-08-24
📖 6 min read🧠 Deep dive

Original authors: Bernhard Hilpert, Kim Baraka, Joost Broekens

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a robot trying to learn a new skill, like cleaning a room or moving a box, while a human watches and tries to help. In the world of robotics, this interaction is called human-interactive learning. The goal is for the robot to get better by listening to the human's feedback. However, a persistent problem has been that humans and robots often speak different languages. When a human sees a robot make a mistake, they might want to explain why it was wrong or suggest a better way to think about the task. But the robot's computer brain is usually designed to only understand simple rewards or penalties, like a score that goes up or down. This mismatch means that when humans try to teach, they often end up forcing their natural teaching style into a rigid box that doesn't fit, leading to confusion and poor results. Researchers have long suspected that to fix this, they need to understand how humans actually think when they teach, rather than just guessing what signals to send.

A team of researchers set out to uncover this hidden logic by stepping back from the usual high-pressure experiments where humans are forced to react instantly to a robot's errors. Instead, they created a calm, observational study to see what people would naturally do if they were simply watching a robot learn without the pressure of having to fix it in the moment. They recruited 34 adults and showed them videos of two different robots learning tasks. One robot was a digital agent navigating a grid to clean up virtual dirt, while the other was a robotic arm trying to push a medication box to a target. The participants watched these robots at three different stages: early in the learning process, in the middle, and near the end. At each stage, the researchers asked a simple, open-ended question: "If you could, how would you help the robot in this video to learn the task better?" The participants were free to answer in any way they wanted, with no restrictions on what kind of help they could offer.

The researchers analyzed hundreds of these responses and discovered that human teaching is not a single, simple action. Instead, it is a complex, layered process that the team named the TOSS framework, standing for Triggers, Objectives, Signals, and Strategies. The first layer, Triggers, revealed that humans do not just wait for a robot to fail. They constantly evaluate the robot's behavior against three different standards. They look at whether a mistake is a consistent pattern or just a one-time glitch. They judge the robot either against its own past performance or against the absolute rules of the task. Finally, they decide if the behavior aligns with their goals or violates them. This means a human teacher is constantly running a sophisticated internal checklist, deciding not just if the robot is wrong, but how and why it is wrong.

Once a human decides a robot needs help, they have a specific goal in mind, which the researchers call Objectives. These goals fall into three distinct categories. Sometimes, the teacher wants to fix the robot's immediate actions, such as making its movements smoother or more precise. Other times, the goal is to ensure the robot completes a specific part of the task, like picking up an object correctly before moving it. Surprisingly, the study found that humans also have a third type of goal: they want to teach the robot about the world itself. Participants often expressed a desire to give the robot a map, explain where objects are located, or clarify the purpose of the task. This suggests that humans intuitively believe robots have an internal mind that can understand facts and concepts, not just react to rewards.

To communicate these goals, humans naturally choose from a variety of Signals. They might act as a judge, giving praise or criticism. They might act as a tutor, offering specific corrections like "slow down" or "turn the other way." They might act as a demonstrator, showing exactly what to do. They might act as a source of information, simply telling the robot facts about the environment. Or, they might choose to say nothing at all, a signal called Withholding, which is a deliberate decision to let the robot figure things out on its own to test its learning. The study showed that these signals are not random; they are carefully chosen tools to bridge the gap between what the robot is doing and what the human wants it to achieve.

Perhaps the most revealing finding was the role of Strategies, which represents the overall role the human teacher adopts. The researchers found that people spontaneously shift between three main identities. Sometimes they act as a Coach, giving real-time feedback and nudges. Other times, they act as an Engineer, deciding that the robot's hardware or software is fundamentally flawed and suggesting new sensors or better algorithms. In other cases, they act as a Designer, changing the environment itself to make the task easier, such as removing obstacles or rearranging the room. This discovery challenges the current way robots are built. Most systems assume a human will only ever be a Coach, offering simple feedback. But this study shows that when people are free to think naturally, they often want to redesign the robot or the task entirely.

The researchers concluded that for robots to learn effectively from humans, the interaction needs to change. Current systems force humans into a narrow role, ignoring their natural tendency to act as engineers or designers. By understanding the TOSS framework, future robots could be designed to recognize when a human is shifting from a Coach to an Engineer. This would allow the robot to adapt its learning process accordingly, perhaps by accepting a new map or a changed environment instead of just waiting for a simple command. The study does not claim to have solved the problem of robot learning, but it provides a clear, detailed map of how humans actually think when they teach. It suggests that the key to better human-robot collaboration lies in building systems that respect the full complexity of human intuition, allowing people to teach in the rich, multi-layered way they naturally do.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →