← Latest papers
📄 health informatics

Comparing Human and Large Language Model Responses to Patients Online Questions: Towards Multi-dimensional Patient-centered Support

This study empirically compares large language models and peer responses to patients' online questions about laboratory test results, finding that while LLMs excel at providing clear, structured medical explanations, peers offer more personalized emotional support, suggesting that LLMs could effectively complement peer communities if they improve their emotional depth, reasoning transparency, and alignment with community norms.

Original authors: Hussein, M. A., Doshi, R., He, L., Reynolds, T.

Published 2026-07-17
📖 7 min read🧠 Deep dive

Original authors: Hussein, M. A., Doshi, R., He, L., Reynolds, T.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you're standing in a vast, bustling digital town square called an Online Health Community. It's a place where people who are sick, worried, or just curious about their bodies gather to ask questions and share stories. In this town, there are two main types of helpers. First, there are the Peers: regular people like you and me, who have been through similar health scares. They offer a warm hug, a shared story, and a "I've been there" kind of comfort. Second, there are the Large Language Models (LLMs). Think of these as incredibly smart, super-fast robots that have read almost every book and website ever written. They can explain complex medical terms like a professor and organize information better than a librarian.

The big question scientists are asking right now is: When someone posts a confusing lab test result online, who is better at helping them? Is it the empathetic human who understands the fear, or the super-smart robot that knows the facts? This study dives into that exact scenario. It looks at how these two very different helpers react when people post about their lab results, checking to see who is easier to read, who offers better emotional support, and who gives the most personalized advice. The goal isn't to replace the humans with robots, but to figure out if these robots can be useful sidekicks in the town square without causing confusion or hurt feelings.


The Great Showdown: Humans vs. Robots in the Health Chat Room

The researchers set up a fascinating experiment. They took 122 real questions from an online health community where people had posted their lab test results (like blood work) and asked for help. Then, they fed these exact same questions into four different AI "brains" (GPT-3.5, GPT-4o, DeepSeek-R1, and Mistral-7B). They compared the robots' answers to the 519 replies written by actual human users. Here is what they found, broken down into four key categories.

1. The Length and Reading Level: The Robot's Ramble

If you asked a human for help, they might send a quick, friendly text. If you asked a robot, you might get a novel. The study found that human replies were short, averaging about 5.09 sentences. The robots, however, were much chatty. GPT-3.5 averaged 9.26 sentences, GPT-4o wrote 15.83 sentences, DeepSeek-R1 went on for 18.92 sentences, and Mistral-7B wrote 13.06 sentences.

It wasn't just that the robots talked more; they also used bigger words. The researchers measured "readability" using two tools (ARI and SMOG). Human comments were written at a level a high school student could easily understand. The robots, especially GPT-3.5, wrote at a slightly higher, more complex level. While the robots were clear and structured, they were sometimes too wordy and used vocabulary that might be a bit tough for someone just trying to understand a simple test result.

2. Emotional Support: The Hug vs. The Pep Talk

This is where things got really interesting. You might think robots are cold and humans are warm, but the data showed a surprise twist.

  • The Humans: About 38.4% of human replies offered very little emotional support. They were often just giving facts or vague encouragement. However, when humans did get emotional, they were incredibly diverse. They offered comfort (44.3%), reassurance (29%), and empathy (12.9%). They shared personal stories like, "I know exactly how you feel because I went through this too," which made people feel less alone.
  • The Robots: The robots were actually more likely to offer emotional support than the average human. About 65% to 71% of robot replies were rated as "high" in emotional support. But here's the catch: their support was very one-note. They mostly relied on encouragement (ranging from 42.9% to 82% depending on the robot). They would say things like, "You can do this!" or "Stay strong!" but they rarely offered the deep, specific comfort or shared lived experience that humans did.

The Metaphor: Imagine you are crying because you're scared of a medical result. A human might sit next to you, cry with you, and say, "I remember when I got this result, I was terrified too, but here is what happened next." A robot might stand up, hand you a tissue, and say, "I understand you are scared. It is a tough time, but you are strong and will get through this." Both are helpful, but the robot's hug feels a bit more like a pep talk from a coach than a hug from a friend.

3. Personalization: The Mirror vs. The Storyteller

Who paid more attention to the specific details of the person's story? Surprisingly, the robots won this round.

  • The Robots: Over 90% of the robot replies were "highly personalized." They acted like perfect mirrors. If you told them your TSH level was 47.15 and you were tired, they would repeat those exact numbers back to you and organize them neatly. They were great at taking your messy story and turning it into a structured list.
  • The Humans: Only 64% of human replies were highly personalized. Humans often gave generic advice like "I hope you feel better" without looking closely at the specific numbers you posted. However, when humans were personalized, they did it differently. They didn't just mirror your facts; they added their own life experience. They might say, "Your husband drinking alcohol with Hepatitis C is dangerous because I saw my brother make that mistake."

The Metaphor: The robot is like a super-organized notepad that writes down everything you said and adds a neat summary. The human is like a friend who listens to your story and then tells you a story about their cousin who had the same problem. The robot is better at the "notepad" part; the human is better at the "storytelling" part.

4. Transparency: The "Why" vs. The "What"

Transparency means explaining how you got your answer. Did the helper show their work?

  • The Humans: About 51.6% of human replies were highly transparent. They would explain their reasoning step-by-step, like, "I think this is the case because my doctor told me X, and I read Y." Even when they were wrong, they were honest about where they got their info.
  • The Robots: Only about 28% to 32% of robot replies were highly transparent. While they were polite and admitted they were AI, they often gave answers without showing the deep reasoning behind them. They would say, "It's important to talk to your doctor," but they wouldn't always explain why that specific advice was given in that specific context. They were often vague to stay safe.

The Final Verdict: Teamwork Makes the Dream Work

So, who wins? The paper suggests that neither side wins alone. The robots are amazing at organizing information, explaining medical terms clearly, and providing a structured, supportive response quickly. But they lack the deep, messy, real-world empathy and the "lived experience" that humans bring. Humans are great at emotional connection and sharing personal stories, but they can be inconsistent, sometimes vague, and not always the best at organizing complex data.

The authors conclude that we shouldn't try to replace the human town square with robots. Instead, the robots should be the sidekicks. Imagine a system where a robot jumps in first to say, "Here is a clear explanation of your lab results, and here is a summary of what they mean." Then, the human community steps in to say, "I've been there, and here is how I handled it."

The paper warns that we aren't ready to just let robots run the show yet. They need to get better at understanding emotions, showing their work more clearly, and respecting the rules of the online communities. But if we use them carefully—as helpers who fill in the gaps rather than taking over—they could be a powerful tool to help people feel less alone and more informed.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →