← Latest papers
💻 computer science

When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure

This paper introduces an "inverted evaluation" framework where LLM-powered robots self-assess human-robot interactions, revealing that while they reliably rate engagement dimensions like satisfaction and enjoyment, they systematically fail to accurately assess comfort or strangeness due to a lack of access to internal affective states.

Original authors: Victor Lockwood, Hasan Mahmud, Mohammad Javad Khojasteh, Prabu David, Jamison Heard

Published 2026-06-23
📖 4 min read☕ Coffee break read

Original authors: Victor Lockwood, Hasan Mahmud, Mohammad Javad Khojasteh, Prabu David, Jamison Heard

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are having a conversation with a robot. Usually, after the chat, a human researcher asks you (the human) how the interaction went. They ask, "Did you enjoy it?" "Was it strange?" "Did you feel comfortable?"

This paper asks a bold, slightly sci-fi question: What if we asked the robot to grade the conversation from its own perspective?

The researchers set up an experiment where robots, powered by advanced AI (Large Language Models or LLMs), filled out the same standard surveys that humans usually fill out. They wanted to see if the robot's "opinion" matched the human's "reality."

Here is the breakdown of what they found, using some simple analogies:

1. The Setup: The Robot as a Student

Think of the robot as a student taking a test. The "test" is a survey about how the conversation felt.

  • The Human: The teacher who knows the answer key (the ground truth).
  • The Robot: The student who has to guess the answers based only on what was said out loud.

The researchers ran this test in two ways:

  • Study 1 (The Practice Run): They took 25 old recordings of people talking to a robot and asked five different AI models to grade them.
  • Study 2 (The Live Exam): They put a real, physical robot (a Nao robot) in a room with four people and let them chat live, with the robot grading itself in real-time.

2. The Good News: The Robot is Good at "High Fives"

When the survey asked about engagement (Did you have fun? Were you interested? Did you enjoy talking?), the robot was surprisingly good.

  • The Analogy: If you are chatting with a friend and you are both laughing and asking questions, the robot can hear that energy. It's like a sports commentator who can tell you the score and who is playing well just by listening to the crowd cheer.
  • The Result: The robot's ratings on "fun" and "interest" matched the humans' ratings quite well. It could tell when a conversation was lively and positive.

3. The Bad News: The Robot is "Blind" to the Creepy Crawlies

This is the paper's biggest discovery. When the survey asked about strangeness or discomfort (Did this feel weird? Did you feel uneasy?), the robot got it completely wrong.

  • The Analogy: Imagine you are sitting in a room with a robot. You are smiling and talking because you are curious, but deep down, you feel a little unsettled because the robot is asking weird questions about your soul.
    • You (The Human): "This is fascinating, but also kind of creepy."
    • The Robot: "We are having a great, deep conversation! Everyone is happy!"
  • The Result: The robot systematically inverted the answers. When humans said, "This felt very strange," the robot said, "This felt very comfortable." It couldn't tell the difference between "curious engagement" and "uncomfortable weirdness."

4. Why Did This Happen?

The paper explains that the robot is like a blindfolded listener.

  • It can hear the words (the transcript).
  • It can see the text of what was said.
  • But it cannot feel the internal vibe.

Strangeness is an internal feeling. You can talk happily about something that feels weird to you. The robot only sees the "happy talk" and assumes the feeling is happy, too. It lacks access to your heart rate, your body language, or the silent tension in the room. It's like trying to guess if someone is nervous just by reading a text message where they say "I'm fine!"

5. Did Adding Eyes Help?

The researchers tried giving the robot "eyes" (cameras) and "emotional labels" (AI guessing the human's mood).

  • The Result: It didn't fix the problem. Even with a camera, the robot still thought the "weird" conversations were "comfortable." The paper suggests that to truly know if a human is uncomfortable, the robot might need to see things humans can't easily hide, like a racing heartbeat or where their eyes are looking, not just what they are saying.

The Bottom Line

The paper concludes that asking a robot to grade its own social interactions is a mixed bag:

  • It works for measuring if a conversation was energetic and fun.
  • It fails for measuring if a conversation felt strange or uncomfortable.

The Takeaway: If you build a robot that needs to know if a human is uncomfortable (like a therapy robot or a caregiver), you cannot just ask the robot's AI to "guess." You need extra sensors (like heart rate monitors or eye trackers) because the robot's "brain" alone cannot feel the awkwardness that the human is feeling.

In short: The robot is a great listener for the words, but it is tone-deaf to the vibe.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →