← Latest papers
💻 computer science

LLMs for Survey Text Analysis – A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis

This study demonstrates that GPT-5.4 can approximate human performance in inductive content analysis of survey data, achieving coding and theme generation agreement levels comparable to human internal consistency, thereby validating its potential as a scalable support tool for qualitative research.

Original authors: Leonardo Bergmann, Renata Gheorghiu, Ana Gvritishvili, Alex Mican, Chris Stewart, Topias Tolonen-Weckström

Published 2026-09-16
📖 5 min read🧠 Deep dive

Original authors: Leonardo Bergmann, Renata Gheorghiu, Ana Gvritishvili, Alex Mican, Chris Stewart, Topias Tolonen-Weckström

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Qualitative research is the art of making sense of words. When scientists, sociologists, or policy makers need to understand the thoughts, feelings, and experiences of people, they often ask open-ended questions. Instead of ticking a box, a person writes a paragraph about their life, their struggles, or their hopes. To turn thousands of these paragraphs into useful knowledge, researchers must read every single one, find the repeating ideas, and group them into categories. This process, known as content analysis, is the backbone of understanding human behavior, but it is incredibly slow and demanding. For decades, the only way to do this was with human minds, reading line by line. Now, a new kind of tool has entered the room: large language models. These are artificial intelligence systems trained on vast amounts of text that can read, understand, and summarize language. The big question for the scientific community is not just whether these machines can read, but whether they can think like a human researcher when the task requires finding hidden patterns without a pre-written rulebook.

A team of researchers from universities across Europe set out to answer this question with a direct comparison. They took a real-world dataset from a survey of 2,800 PhD students across Europe, where 10% of the responses were randomly drawn to serve as the data basis for the study. From this sample, a total of 903 responses across six specific open-ended questions were analyzed in depth. On one side of the experiment, five human researchers read these answers and performed what is called inductive content analysis. This means they did not start with a list of categories to force the answers into. Instead, they read the text, identified the core ideas, created their own labels for those ideas, and then grouped those labels into broader themes. It is a creative and interpretive process, relying on human judgment to decide what matters. On the other side, the researchers used a sophisticated language model to perform the exact same task on the same text. The machine was given the same instructions to read the answers, find the main ideas, and group them into themes, all without being told what to look for in advance.

The researchers then compared the two sets of results to see how closely the machine matched the humans. They did not look for a perfect match, because even different human experts often disagree on how to label a specific sentence. Instead, they measured the overall alignment of the groups. The results showed a moderate level of agreement. When looking at the specific labels the researchers and the machine created for the text, they agreed about 61 percent of the time. When looking at the broader themes they built from those labels, the agreement was slightly lower, at 54 percent. To put this in perspective, the researchers also checked how much the five human experts agreed with each other. The humans agreed with each other about 68 percent of the time on the labels and themes. This means the gap between the human and the machine was not huge; the machine performed almost as consistently as a human expert does with another human expert.

However, the story is not uniform across all topics. The agreement between the humans and the machine varied significantly depending on what the students were writing about. For questions regarding financial situations, the machine struggled more, showing lower alignment with the human researchers. For questions about personal experiences or general feedback, the alignment was much stronger. The study suggests that the reliability of the machine depends heavily on the nature of the data and the clarity of the text. When the text was ambiguous or the internal logic of the machine's own grouping was shaky, the agreement with humans dropped. The researchers found that whenever the machine was inconsistent with itself, it was also inconsistent with the humans, and the same was true for the human team. This indicates that the difficulty of the specific topic plays a major role in how well the tool works.

The human coders who participated in the study were also asked for their impressions of the machine's work. They reported that the machine's categorizations reflected their own thinking and that they would trust the tool to help analyze future data. The researchers concluded that these artificial intelligence systems are not ready to replace human judgment entirely, especially for the complex, high-level work of building theories. However, the findings suggest that these tools are highly effective at handling the early, labor-intensive stages of sorting through text. They can do the heavy lifting of reading thousands of responses and finding the initial patterns, allowing human researchers to focus on the deeper interpretation and validation of those patterns. The study does not claim that the machine is perfect or that it has solved the problem of text analysis, but it does suggest that it is a powerful, scalable partner for researchers who need to make sense of large volumes of human stories.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →