← Latest papers
📄 medicine

Quality and Readability of Artificial Intelligence Chatbot Responses to Patient Questions About Exocrine Pancreatic Insufficiency: A Comparative Study

This comparative study found that while four major AI chatbots generated well-structured and broadly useful information about Exocrine Pancreatic Insufficiency, their responses lacked source transparency, exhibited poor treatment information quality, and significantly exceeded the recommended sixth-grade reading level, necessitating professional verification before use.

Original authors: Donghong Wang, Baihui Chen, Jialin Wang, Yalu Chen, Xian Deng, Xiaoyi Huang, Qing Chen

Published 2026-08-10
📖 4 min read☕ Coffee break read

Original authors: Donghong Wang, Baihui Chen, Jialin Wang, Yalu Chen, Xian Deng, Xiaoyi Huang, Qing Chen

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet as a giant, bustling library where you can ask any question and get an answer instantly. In recent years, a special kind of librarian has arrived: the Artificial Intelligence (AI) chatbot. These aren't just search engines that list links; they are like super-smart, chatty robots that can write full stories, explain complex ideas, and even pretend to be doctors or teachers. But here's the catch: just because a robot can talk fluently doesn't mean it's telling the truth, or that it's speaking a language you can actually understand.

This paper dives into a specific corner of health science called "Exocrine Pancreatic Insufficiency," or EPI for short. Think of your pancreas as a tiny factory inside your belly that pumps out special soap-like chemicals (enzymes) to help you digest food. When this factory breaks down, you can't absorb nutrients properly, leading to tummy trouble and weight loss. Patients with this condition often turn to the internet for answers, hoping to understand their diagnosis and how to manage it. The big question for this study is: If a patient asks an AI robot about EPI, will the robot give them clear, trustworthy, and easy-to-read advice, or will it just sound smart while confusing everyone?

The Great Robot Read-Off

To find out, the researchers set up a digital showdown. They gathered 20 burning questions that real patients might ask about EPI, ranging from "What is this disease?" to "What should I eat?" and "How do I take my medicine?" They then fed these questions into four of the most popular AI chatbots available: ChatGPT, Copilot, Gemini, and Perplexity. The goal was to see which robot was the best teacher.

The team acted like strict judges, using four different scorecards to grade the robots' answers. One scorecard checked if the advice was safe and reliable (DISCERN), another looked at how well-organized the information was (EQIP), a third checked if the robots admitted where they got their facts (JAMA), and a fourth gave a general "thumbs up or down" for usefulness (GQS). They also ran a "reading level test" to see if a sixth-grader could understand the answers, which is the gold standard for patient education.

The Verdict: Good Structure, Bad Reading Level

Here is the twist: The robots were surprisingly similar. The study found no significant difference between the four AI models. Whether you asked ChatGPT, Copilot, Gemini, or Perplexity, they all performed about the same.

When it came to organization and general usefulness, the robots did a decent job. They gave answers that looked neat, flowed well, and seemed helpful on the surface. However, when the judges looked deeper, the cracks appeared. The robots scored poorly on "treatment information quality," meaning they didn't always explain the risks, benefits, or uncertainties of treatments clearly. Even worse, they completely failed the "source transparency" test. None of the robots told the user where they got their information, who wrote it, or when it was last updated. It was like getting a recipe from a chef who refuses to tell you where they bought the ingredients or if the recipe is from 1990.

The biggest problem, however, was the language. The researchers wanted to know if the answers were simple enough for a sixth-grader to read. The answer was a resounding no. Every single robot wrote at a level far too complex for the average patient. The reading scores were equivalent to high school or even college-level text. While Perplexity and Copilot were slightly easier to read than the others, they still didn't meet the basic standard. The robots used long, complicated sentences and big medical words without explaining them, making the information hard to digest for the very people who needed it most.

What This Means for You

The study concludes that while these AI chatbots are great at sounding professional and organizing information, they are not ready to be your sole source of medical advice. They are like a very well-dressed tour guide who knows the map but speaks in riddles and never shows you the source of their directions.

The authors suggest that patients can use these robots to get a basic idea of what a disease is or to prepare questions for a real doctor. However, they should never use the robots to decide on their own treatment, change their medication, or interpret test results without a professional checking the work first. The robots need to get better at citing their sources, updating their info, and speaking plain English. Until then, they are best used as a supplement to, not a replacement for, a real human doctor's advice.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →