← Latest papers
🤖 AI

Why did My Robot Just Change Personality? Prompting Guidelines for a Grounded Robot Persona in LLM-Based HRI

This paper addresses the risks of unpredictable robot behaviors in LLM-based human-robot interaction by proposing a structured eight-component prompt design framework and ethical guidelines, grounded in expert feedback, to ensure transparent, safe, and adaptable robot personas.

Original authors: Ashita Ashok, Franziska Babel, Patrick Holthaus, Rucha Khot, Karla Bransky, Fethiye Irmak Dogan, Karsten Berns, Silvia Rossi, Minha Lee, Guy Laban

Published 2026-08-28
📖 6 min read🧠 Deep dive

Original authors: Ashita Ashok, Franziska Babel, Patrick Holthaus, Rucha Khot, Karla Bransky, Fethiye Irmak Dogan, Karsten Berns, Silvia Rossi, Minha Lee, Guy Laban

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a social robot, a machine built to chat, help, and perhaps even comfort a human. To make this conversation feel natural, engineers often connect the robot to a large language model, a type of artificial intelligence trained on vast amounts of text to understand and generate human speech. This technology allows the robot to answer questions, tell stories, and react to emotions with a fluency that was impossible just a few years ago. However, there is a hidden gap between the robot's physical reality and the words it speaks. The machine might not have eyes to see a room, hands to pick up an object, or access to the internet, yet the language model it uses knows about these things because it has read about them. Without careful instructions, the robot might confidently claim it can see you or fetch a cup, creating a confusing and potentially misleading experience for the person talking to it. This disconnect is the central puzzle facing researchers who want these machines to be safe, honest, and truly helpful partners.

A team of researchers from universities across Europe, Australia, and Israel recently tackled this problem by asking a simple but critical question: how do we tell a robot exactly who it is and what it can do before it starts talking? Their work, presented at a major robotics conference, moves beyond the idea that the robot's personality is just a minor setting to be tweaked. Instead, they argue that the instructions given to the artificial intelligence—the "prompt"—are the most important design element for shaping a robot's behavior. To solve this, the team developed a structured framework that breaks down the robot's identity into eight specific parts. These parts include defining the robot's role, clearly stating its physical limits, explaining how it handles mistakes, and setting strict ethical boundaries to prevent it from lying or causing harm.

The researchers did not just invent this list in a quiet office; they built it by looking at what other scientists were already doing and then testing those ideas with experts. First, they reviewed recent studies where robots used large language models to talk to people. They found that most researchers were only telling the robot its name and its main job, like "be a tutor" or "be a companion." They were rarely telling the robot what it could not do, such as admitting it has no memory of past conversations or that it cannot see the user's face. This omission meant that robots often acted as if they had abilities they did not possess, leading to confusion and broken trust.

To understand why this matters, the team gathered twenty-seven experts in human-robot interaction for a workshop. These were seasoned researchers and practitioners from diverse backgrounds, including psychology, ethics, and computer science. They discussed how a robot's personality should work over time and what happens when a robot makes a mistake. The experts agreed on a crucial point: a robot's personality is not just about the words it says. If a robot claims to be a caring friend but then forgets a conversation the user had five minutes ago, the user feels deceived. The experts noted that subtle cues of personality are often missed by people, and that a robot's identity must be consistent with its actual physical abilities. They also raised serious concerns about safety, noting that if a robot lies about its capabilities, it could lead to dangerous situations, especially for vulnerable users like children or the elderly.

Based on these discussions and their review of existing work, the team proposed a new way to write the instructions for these robots. Their framework requires designers to explicitly write down eight things before the robot ever speaks. First, they must define the Identity, which is the robot's specific character and social role. Second, they must set Capability Boundaries, a clear list of what the robot can and cannot perceive or do, such as stating it has no internet access or cannot move on its own. Third, Transparency instructions must tell the robot how to admit it is a machine and what its limits are. Fourth, the Task defines the specific goal of the conversation. Fifth, an Expectation and Failure Protocol explains how the robot should react when it does not know an answer or when the conversation goes wrong. Sixth, Privacy rules dictate how the robot handles personal data. Seventh, User Adaptation ensures the robot changes its tone for different people, such as speaking simply to a child. Finally, Ethical Red Lines set hard limits to prevent the robot from saying anything harmful, biased, or illegal.

The researchers illustrated this idea by providing two complete prompt examples using all eight modular design components as a proof-of-concept template in a separate project. They argue that when these components are included, the robot's behavior becomes much more predictable and honest. The experts at the workshop confirmed that without these specific instructions, a robot's personality can feel unstable or deceptive. They emphasized that a robot cannot simply be told to "be friendly"; it must be told exactly how to be friendly within the limits of its hardware. For instance, if a robot is designed to help an elderly person, it must be programmed to know it cannot physically lift them, even if the language model suggests it could.

This work suggests that the future of social robots depends less on making the language sound more human and more on making the robot's limitations clear. The team argues that treating the prompt as a serious design tool, rather than a simple coding step, is essential for building trust. By forcing designers to write down exactly what the robot is and what it is not, the field can move away from robots that hallucinate capabilities and toward machines that are reliable, transparent, and safe. The study does not claim to have solved every problem in robotics, but it provides a concrete, structured way to ensure that when a robot speaks, its words match its reality. Future work should explore the proposed prompt guidelines implemented in different robot embodiments to evaluate the perception of the robot identity.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →