Artificial Intelligence in Surgical Qualitative Research: A Comparison of Human and AI-Assisted Thematic Analysis
This study compares human and AI-assisted thematic analysis in surgical qualitative research, finding that while both methods identify similar overarching themes, human-generated codebooks yield higher inter-coder reliability due to clearer thematic boundaries, suggesting a hybrid approach combining AI efficiency with human refinement may be optimal.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a giant box of 200 handwritten notes from surgeons. Each note answers one question: "What is the hardest thing about doing a specific surgery (LCBDE) at your hospital?"
Your goal is to read all these notes, sort them into piles based on the problems mentioned, and give each pile a label. This process is called Thematic Analysis. Usually, humans do this by reading, arguing, and refining the labels until they agree on what each pile means.
This paper asks a simple question: Can a robot (Artificial Intelligence) do this sorting job just as well as a human?
Here is the story of how they tested it, using some simple analogies.
The Two Teams
The researchers set up a race between two teams to sort the same 200 notes:
- The Human Team: Two doctors read the notes. They started with no labels. As they read, they created their own list of categories (like "Not enough money," "Not enough time," "Bad equipment"). They kept refining these labels until they were very clear and distinct. This is their Human Codebook.
- The AI Team: They fed all 200 notes into a powerful AI (ChatGPT) in one big batch. They told the AI, "Look at these notes and make your own list of categories." The AI did this instantly without any human help or examples. This is the AI Codebook.
The Results: Who Sorted Better?
Both teams managed to find the same big picture problems. Whether it was a human or a robot reading the notes, they both agreed that the main issues were:
- Equipment problems
- Money/Cost issues
- Time constraints
- Lack of support from bosses
- Not doing the surgery often enough (Low Volume)
However, there was a difference in how they sorted the notes.
The "Fuzzy Pile" Problem
Think of the Human Team's labels like clear, distinct boxes. If a note said, "We don't have the right tools," it went into the "Equipment" box. If it said, "We don't have time," it went into the "Time" box. The boxes didn't overlap much.
The AI Team's labels were a bit more like fuzzy, overlapping circles. The AI created some very broad categories. For example, it might have a label called "Workflow and Culture." A single note about "The hospital prefers a different surgery" could fit into three different AI categories at the same time.
Because the AI's categories were so broad and overlapping, the two human coders who tried to use the AI's list to sort the notes got confused more often. They didn't always agree on which "fuzzy" pile a note belonged to.
The Scorecard
The researchers measured how often the two coders agreed with each other (Inter-Coder Reliability).
- Human Team: They agreed 95.1% of the time. Their "agreement score" (Kappa) was 0.75.
- AI Team: They agreed 92.3% of the time. Their "agreement score" was 0.64.
While the AI was still pretty good, the humans were slightly more consistent. The paper notes that the difference was statistically significant, meaning it wasn't just a fluke; the human labels were just a bit clearer.
The "Hybrid" Solution
The paper concludes that the AI is like a fast, helpful assistant who can quickly look at a messy room and say, "Okay, here are the main types of clutter: clothes, books, and dishes."
But the AI might not be perfect at defining exactly where a "sock" goes if it's mixed with a "towel." The humans are needed to come in afterward and say, "Actually, let's make a specific rule for socks so everyone knows exactly where they go."
The Takeaway:
You can use AI to do the heavy lifting of finding the main themes quickly. But to make sure the rules are clear and everyone sorts the notes the same way, a human needs to step in and refine the definitions. The best approach is a hybrid: let the AI do the first pass, then let humans polish the labels.
What the Paper Doesn't Say
It is important to note what this study did not claim:
- It did not say AI is ready to replace human researchers entirely.
- It did not test the AI on long, complex interviews (only short survey answers).
- It did not test different versions of AI or different ways of asking the AI questions (it used a "zero-shot" method, meaning it asked the AI to do it with no prior training or examples).
- It did not claim the AI was "wrong" about the themes; the AI found the right big ideas, it just had trouble drawing the lines between them as clearly as humans did.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.