Large Language Models for Pulmonary Embolism Health Education: A Four- Model Comparison of Reliability, Quality, and Readability
This study compares four large language models on pulmonary embolism health education, finding that while Copilot and Perplexity demonstrated superior reliability and sourcing, none of the models met recommended readability standards for patient education, highlighting the need for professional oversight.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet as a giant, bustling library where anyone can walk in and ask a question. In the past, if you asked about a scary medical condition, you might get a stack of old books, a confusing newspaper article, or a wild rumor from a stranger. But recently, a new kind of librarian has arrived: the Large Language Model (LLM). Think of these as super-smart, tireless robots that have read almost everything written on the internet and can chat with you in plain English. They are great at writing stories, solving math problems, and even answering health questions.
But here is the tricky part: when it comes to serious health issues, being a good storyteller isn't enough. You need a librarian who knows the facts, admits when they aren't sure, and explains things simply enough for a tired, worried person to understand. One such scary condition is a Pulmonary Embolism (PE). It's like a traffic jam in the lungs caused by a blood clot, and it can be very dangerous. Because it's so serious, people often rush to Google to find answers. This paper asks a big question: If you ask four different AI librarians about PE, will they give you the same good advice? Or will some sound confident but be wrong, or give answers that are too hard to read?
The Great AI Showdown: Who's the Best Health Librarian?
In this study, researchers from Xinfeng County People's Hospital decided to put four of the most popular AI chatbots to the test. They picked ChatGPT, Gemini, Microsoft Copilot, and Perplexity AI. To make it a fair fight, they didn't ask the robots weird, made-up questions. Instead, they looked at what real people were actually searching for on Google over a five-year period. They found the 31 most common questions people ask about Pulmonary Embolism—things like "What are the symptoms?", "How is it treated?", and "Is it dangerous during pregnancy?"
Then, they fed these exact 31 questions into each of the four AI models. Since there were four models and 31 questions, the researchers ended up with 124 different answers to grade. They acted like strict teachers, using four different "report cards" to check the answers:
- Reliability: Did the AI admit what it knew and what it didn't? Did it show its homework (sources)?
- Quality: Was the advice balanced and useful?
- Readability: Was the language simple enough for a 6th grader to understand? (Doctors usually recommend health info be this simple so everyone can get it).
- Overall Flow: Did it read smoothly?
The Results: The "Good" News and the "Bad" News
Here is the plot twist: All four AI models answered every single question. None of them crashed or refused to talk. They all sounded confident and wrote in clear, flowing sentences. If you just read the first paragraph, they all seemed like helpful experts.
However, when the researchers looked closer, the differences were huge.
The "Source" Score (JAMA Benchmark):
Imagine an essay where you get points for listing your sources. Two models, Copilot and Perplexity, were the only ones that bothered to show their work. They gave citations (links to where they got their info). The other two, ChatGPT and Gemini, mostly just gave answers without showing where they came from. In fact, ChatGPT and Gemini got a score of 0 for showing sources, while Copilot and Perplexity scored slightly higher (around 1 out of 4). This means if you wanted to check if the AI was telling the truth, you could only really do that with the first two.
The "Reliability" Score (DISCERN):
This score checks if the advice is balanced and safe.
- Copilot and Perplexity scored a median of 41.
- ChatGPT scored 35.
- Gemini scored 33.
On a scale where 63 is "excellent" and 16 is "very poor," all four models landed in the "poor" to "fair" range. Even the "winners" (Copilot and Perplexity) didn't reach the "good" level. They were okay, but not great.
The "Readability" Problem:
This is where the plot gets a bit frustrating. The American Medical Association says health advice should be written at a 6th-grade reading level so that anyone, regardless of their education, can understand it.
- The researchers checked the answers using six different math formulas to measure how hard the text was to read.
- None of the four models passed the test.
- The easiest answers were still written at a 9th-grade level (ChatGPT).
- The hardest answers were written at a 16th-grade level (Copilot), which is like college-level reading!
- The "Flesch Reading Ease" score (where higher is easier) was around 35 to 48 for all of them. To pass, they needed a score of 80 or higher.
Basically, the AI robots were writing in a language that is too complicated for the average person to read quickly, especially when they are scared and worried about a medical emergency.
The Verdict: Don't Let the Robot Drive the Car
The study found that while these AI models are great at sounding smart and organizing information, they aren't ready to replace a doctor.
- Copilot and Perplexity were the "better" students because they showed their sources and had slightly better structure, but they still wrote in too many big words.
- ChatGPT and Gemini were a bit worse at showing sources, even though they sounded just as smooth.
- Crucially, the study found that no model combined reliable facts, clear sources, AND simple language.
The authors conclude that these AI tools are like a helpful co-pilot, but they should never be the one driving the car. They can help explain what a disease is or what questions to ask a doctor, but they cannot diagnose you, tell you if you are safe, or decide on your treatment. If you use an AI for health info, you must remember that it might be confident but wrong, and it might be too hard to read. The best plan is to use the AI to get a starting point, but then have a real human doctor check the work.
In short: The AI librarians are polite and fast, but they are still learning how to speak "human" and how to prove they aren't making things up. Until they get better, we need to keep a human in the loop.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.