← Latest papers
💻 computer science

Evaluating Human and LLM-Generated Thematic Analysis in HRI for Vulnerable Populations: A Comparative and Ethical Analysis

This paper presents a comparative and ethical analysis of human versus LLM-generated thematic analysis in Human-Robot Interaction research involving vulnerable populations, evaluating their agreement and investigating whether LLMs risk misrepresenting or marginalizing participant experiences.

Original authors: Alva Markelius, Fethiye Irmak Dogan, Julie Bailey, Hatice Gunes

Published 2026-08-25
📖 4 min read☕ Coffee break read

Original authors: Alva Markelius, Fethiye Irmak Dogan, Julie Bailey, Hatice Gunes

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the field of human-robot interaction, researchers often turn to social robots to assist people who face significant challenges in daily life, such as individuals with disabilities or children with chronic illnesses. To understand how these people truly experience the technology, scientists rely on a method called thematic analysis. This is a careful, human process where researchers read through hours of interview transcripts, looking for patterns in what people say. It is not merely counting words; it requires deep empathy, cultural awareness, and the ability to understand the subtle, lived reality of the person speaking. For decades, this work has been considered an inherently human craft, one that demands a reflexive engagement with the data to ensure that the voices of vulnerable participants are heard accurately and respectfully.

Recently, a new tool has entered the laboratory: large language models. These are powerful computer programs trained on vast amounts of text that can read, summarize, and even identify patterns in human language. As these models become more capable, researchers have begun to ask if they can help with the heavy lifting of thematic analysis. Could a computer read the interviews and find the same themes a human would? While early tests in general settings showed promise, a critical question remained unanswered: does this work when the data comes from vulnerable populations? If a computer misinterprets the experience of a disabled person, the error is not just a statistical glitch; it risks silencing the very people the research aims to help.

A team of researchers at the University of Cambridge set out to answer this question by putting human and artificial intelligence side by side. They took interview data from a previous study involving 31 university students with disabilities, including those with autism, specific learning differences, and mental health conditions. These students had interacted with social robots and voice agents designed to help them navigate university life. The researchers asked a human expert, who specializes in disability and education, to analyze the transcripts using the standard, rigorous method. At the same time, they fed the exact same text into a sophisticated language model, instructing it to follow the same steps and find the same themes.

The results showed that the computer was not entirely lost. When the researchers looked at how the human and the model grouped the students' comments, the computer performed significantly better than random chance. It successfully identified that certain topics, like the difficulty of interacting with the robot or the need for the robot to understand disability, were related. In fact, for some straightforward topics, the computer's labels were very similar in meaning to the human's. However, the agreement was far from perfect, and the differences were not random. As the analysis moved from simple summaries to deeper, more abstract ideas, the computer began to drift.

The most significant findings emerged when the researchers examined the nature of the disagreements. They found that the computer often treated the interviews as a surface-level exercise, picking out the most obvious words rather than understanding the deeper meaning. For instance, when a student spoke about the specific fear of "ableism"—prejudice against people with disabilities—the computer often replaced this precise term with the much broader and less specific word "discrimination." While this might seem like a minor swap, it effectively erased the specific reality of the student's experience, flattening a unique struggle into a generic category. In other cases, the computer missed the emotional weight of a comment entirely, focusing instead on the practical utility of the robot, thereby ignoring the student's feelings of anxiety or isolation.

The study suggests that while these artificial intelligence tools can be useful for the early, mechanical stages of research, such as sorting through large amounts of text to find initial patterns, they are not yet ready to replace human judgment in sensitive contexts. The computer struggles with the nuance of power, identity, and lived experience. It tends to generalize individual stories into population-level trends, a habit that can marginalize the very participants the research is meant to support. The researchers conclude that for studies involving vulnerable groups, the human element remains indispensable. The computer can assist, but it cannot be trusted to interpret the heart of the human experience on its own, as doing so risks misrepresenting the voices of those who need to be heard most clearly.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →