← Latest papers
🤖 AI

Can you see how I learn? Human observers' inferences about Reinforcement Learning agents' learning processes

This paper investigates how human observers infer Reinforcement Learning agents' learning processes through a novel observation-based paradigm, identifying four core interpretive themes (Goals, Knowledge, Decision Making, and Learning Mechanisms) that evolve over time to inform the design of more transparent and interpretable Human-Robot Interaction systems.

Original authors: Bernhard Hilpert, Muhan Hou, Kim Baraka, Joost Broekens

Published 2026-08-24
📖 6 min read🧠 Deep dive

Original authors: Bernhard Hilpert, Muhan Hou, Kim Baraka, Joost Broekens

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

When a human teaches a child to ride a bicycle, the lesson is a conversation of glances, wobbles, and encouraging words. The child learns by feeling the balance shift, and the parent learns by watching how the child interprets their guidance. This back-and-forth is the essence of hybrid intelligence: two different minds working together to solve a problem neither could master alone. In the world of robotics, this partnership is called human-interactive robot learning. Here, a person acts as a teacher, guiding a robot through a task using feedback or demonstrations. The hope is that the robot will learn to navigate a room, push a box, or clean a floor by understanding the teacher's signals. However, a significant gap often exists between what the teacher intends and what the robot actually understands. If a human thinks the robot is confused, they might offer a hint that the robot interprets as a reward for a completely different action. To build better partnerships, scientists must first understand how humans make sense of a robot's behavior when they cannot speak to it directly.

A team of researchers set out to uncover exactly how people interpret the learning process of a robot when they are only watching. They were not interested in how a robot learns mathematically, but rather in the story a human observer tells themselves about what the robot is thinking. To do this, they developed a new way to study human perception that removed the complication of active teaching. Instead of asking people to interact with a robot and give feedback, the researchers simply asked them to watch videos of robots learning and then describe what they believed was happening inside the machine's mind. This approach allowed the scientists to see the raw, unfiltered assumptions people make when they try to understand an artificial learner.

In their first study, the researchers interviewed nine people who watched a robot learn to navigate a grid-like environment, avoiding holes and finding an exit. The participants were asked to speak their thoughts aloud as they watched the robot move. The researchers listened for patterns in how these observers explained the robot's actions. They found that people did not just see random movements; they constructed a detailed mental model of the robot's internal life. This model consistently fell into four main categories. First, people assumed the robot had specific goals, such as wanting to reach the exit as quickly as possible or trying to avoid falling. Second, they believed the robot was gathering knowledge, remembering which spots were dangerous and which were safe. Third, observers thought the robot was making decisions, weighing options and choosing paths based on a strategy. Finally, people inferred that the robot had a learning mechanism, a way of updating its understanding based on past mistakes or successes. These findings suggested that humans naturally project a rich, human-like inner life onto the machine to make its behavior understandable.

To confirm these initial ideas and see if they held up in more complex situations, the researchers conducted a second, larger study with thirty-four participants. They expanded the experiment to include two different types of tasks and two different types of learning algorithms. In one scenario, a robot learned to clean a room by navigating around furniture and picking up dirt. In the other, a different robot learned to push a box of medication across a table to a patient who could not reach it. The robots used different methods to learn: one used a simple, table-based method of trial and error, while the other used a more complex, continuous method that mimics how biological brains might process smooth movements. The participants watched videos of these robots at three different stages of learning: early, middle, and late. After each video, they rated how relevant the four themes of goals, knowledge, decision-making, and learning mechanisms were to what they were seeing.

The results of this larger study confirmed that the four themes identified in the first experiment were not just a fluke. Across both the cleaning and the pushing tasks, and regardless of the complex learning method the robot used, people consistently relied on these four categories to make sense of the robot. The researchers found that as the robots got better at their tasks, the way people described them changed. In the early stages, observers often thought the robot lacked knowledge or was simply guessing. As the robot improved, people shifted their view, believing the robot had developed a clear plan and was making deliberate decisions. The study showed that human understanding is not static; it evolves as the robot's behavior changes. People do not just see a machine following code; they see an agent that is learning, remembering, and planning.

The researchers also discovered that the specific details of these assumptions were quite fluid. For instance, when people thought about the robot's goals, they sometimes focused on the final outcome, like reaching the exit, and other times on the quality of the movement, like cleaning efficiently or moving safely. Similarly, when considering how the robot learned, people sometimes imagined it was learning through pure exploration, while at other times they believed it was learning from feedback or by reasoning through the problem. This flexibility suggests that humans have a dynamic framework for interpreting robots, one that they constantly adjust as they watch the robot perform. The study did not find that a person's background, such as whether they had children or pets, changed these fundamental patterns of thought.

This work highlights a critical challenge in the future of human-robot collaboration. Because humans naturally assume that robots have goals, knowledge, and decision-making processes similar to their own, there is a risk that these assumptions will be wrong. If a human believes a robot is learning by reasoning when it is actually just following a simple rule, the human might give feedback that confuses the machine. The researchers suggest that for robots to be truly effective partners, their designers must understand these human assumptions. By recognizing that people will inevitably project a human-like mind onto a robot, engineers can create systems that are more transparent. They can design robots that communicate their actual learning process in a way that aligns with human expectations, reducing the gap between what the robot is doing and what the human thinks it is doing. This alignment is essential for building the kind of hybrid intelligence where humans and machines can truly work together, each understanding the other's role in the shared task.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →