Can People Distinguish Human and AI Agency in Humanoid Teleoperation? A Preliminary Study of Agency Perception
This preliminary study introduces the "Ghost-in-the-Loop" framework to demonstrate that participants often struggle to distinguish between human and AI agency in humanoid teleoperation, with their judgments primarily driven by perceived naturalness and multimodal consistency rather than the actual source of control.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a future where a robot standing in your living room is not just a machine following a script, but a partner in conversation, capable of smiling, gesturing, and responding with the fluidity of a human. This vision relies on teleoperation, a method where a person sits far away and controls the robot's movements and voice in real time, effectively using the machine as a remote body. For years, this has required a human operator to be constantly present, limiting how many people a single person can help or how long a conversation can last. However, the rapid rise of artificial intelligence has introduced a new possibility: systems that can generate speech, facial expressions, and gestures on their own, mimicking a human so closely that the difference becomes hard to spot. This creates a fascinating and slightly unsettling question for anyone interacting with these machines: if a robot is acting, can you tell if a real person is pulling the strings, or if an algorithm is doing the work?
Researchers at Sony Computer Science Laboratories and the University of Tokyo set out to answer this question by building a system they call Ghost-in-the-Loop. This framework allows a humanoid robot to switch seamlessly between being controlled by a human operator and being driven entirely by artificial intelligence, all while looking and sounding exactly the same. The robot, a Unitree G1 model equipped with a screen for a face and cameras to see its surroundings, serves as the stage. When a human is in control, they wear a headset to see what the robot sees and use motion sensors to guide the robot's hands and face. When the system switches to AI mode, the robot listens to the conversation and uses generative models to create its own voice, facial movements, and hand gestures, matching the same visual and auditory style as the human-controlled version. The goal was not to see which was better, but to see if a human observer could even tell the difference.
To test this, the team created a series of short video clips showing the robot in conversation. They recorded eight different interactions, ranging from casual chats to task-oriented discussions, each lasting between five and forty seconds. In half of these clips, a human teleoperator was controlling the robot. In the other half, the AI system generated the behavior. Crucially, the robot's voice, face, and body remained consistent across all clips, so the only variable was the source of the control. The researchers then invited fifty people to watch these videos online. After each clip, the participants were asked to decide whether the robot was being run by a human or a computer, rating their certainty on a scale that ran from negative ten for a definite AI to positive ten for a definite human. They were also asked to explain why they felt that way.
The results suggest that in these brief moments, the line between human and machine is remarkably blurry. The participants struggled significantly to distinguish between the two sources of control. On average, their ratings barely shifted between the human-controlled clips and the AI-controlled ones, with the difference being so small that it was statistically indistinguishable from random guessing. The study indicates that for short interactions, people simply cannot reliably tell if a robot is being piloted by a person or an algorithm. When the researchers looked at what the participants were actually thinking, a clear pattern emerged in their explanations. Those who felt they could make a judgment relied heavily on the overall naturalness of the robot's behavior. They paid close attention to the timing of the conversation, the rhythm of the movements, and whether the voice, the facial expression, and the hand gestures all seemed to work together in harmony.
When the robot's movements felt stiff, the timing felt off, or the different parts of its body seemed out of sync with its voice, participants were more likely to guess it was an AI. Conversely, when the robot moved with a coherent flow, where the gestures and expressions matched the speech perfectly, people were more likely to believe a human was behind the controls. This does not mean the AI was perfect, but rather that the flaws in the AI's performance were subtle enough to be missed in a quick glance. The study suggests that our ability to detect agency depends less on the specific technology used and more on the consistency of the performance. If the robot acts as a unified whole, the brain accepts it as a single agent, regardless of whether that agent is flesh or code.
This preliminary work opens a new chapter in how we might design the robots of the future. It implies that in the near term, we may not need to worry about users constantly questioning whether a robot is real or fake during short exchanges. Instead, the focus can shift to how these systems can blend human and artificial intelligence, perhaps letting a human take over when a conversation gets complex while the AI handles the routine parts. The researchers note that their findings are specific to these short, structured clips and that longer, more complex interactions might reveal different patterns. For now, the study confirms that when a robot speaks, smiles, and moves with enough coherence, the human mind is happy to accept it as a partner, leaving the true nature of its agency a quiet mystery.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.