← Latest papers
📄 medicine

Assessing the Quality, Safety, and Readability of Large Language Model Responses from ChatGPT and Gemini for Knee Osteoarthritis Patient Education

In a comparative evaluation of knee osteoarthritis patient education, ChatGPT demonstrated superior clinician-rated accuracy and completeness compared to Gemini, though both models produced content with readability levels exceeding recommended standards and require professional review before clinical use.

Original authors: habibullah, mubasheera

Published 2026-07-09
📖 4 min read☕ Coffee break read

Original authors: habibullah, mubasheera

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a knee that hurts, and you want to know how to fix it. Instead of calling a doctor, you ask two different super-smart, robot librarians: ChatGPT and Gemini. You ask them 36 different questions about your knee pain, ranging from "What exercises help?" to "Can I take this specific Unani medicine?"

This study is like a report card comparing how well these two robot librarians answered those questions. Here is what the researchers found, explained simply:

The Setup: A Blind Taste Test

The researchers didn't just ask the robots; they had five experienced knee doctors (who had practiced for over 10 years) grade the answers. The doctors didn't know which robot wrote which answer—they were "blindfolded" to the brand. They graded the answers on five things:

  1. Accuracy: Is the fact true?
  2. Completeness: Did they tell the whole story, or just part of it?
  3. Clarity: Was it easy to understand?
  4. Safety: Did they give advice that could hurt you?
  5. Trustworthiness: Did they sound reliable?

They also checked how "hard" the reading level was, like checking if a book is written for a 6th grader or a college professor.

The Results: Who Won?

1. The "Fact-Checker" Score (Accuracy & Completeness)

  • The Winner: ChatGPT.
  • The Analogy: Imagine you are baking a cake. If you ask for a recipe, Gemini gave you a recipe that was mostly correct but missed a few key ingredients (like forgetting to mention you need eggs). ChatGPT gave you a recipe that was not only correct but included all the steps and warnings.
  • The Data: The doctors gave ChatGPT higher scores for being more accurate and more complete. This difference was strong enough that it wasn't just a lucky guess.

2. The "Safety & Trust" Score

  • The Result: It was a tie.
  • The Analogy: Both robots were equally careful not to tell you to do something dangerous (like stopping your medication without asking a doctor). Neither was significantly safer or more trustworthy than the other.

3. The "Reading Level" Score

  • The Winner: Gemini (by a small margin).
  • The Analogy: Think of the reading level like the height of a fence. Both robots built fences that were too tall for the average person to jump over easily. However, Gemini's fence was slightly lower (easier to jump) than ChatGPT's.
  • The Catch: Even though Gemini was "simpler," both robots wrote at a level that is too hard for many patients. They wrote like college students, but health advice should be written like a 6th or 8th grader.

The "Human Factor" Problem

Here is the tricky part of the study: The five doctors grading the answers didn't always agree with each other.

  • The Analogy: Imagine five judges at a talent show. If one judge gives a singer a 4/5 and another gives them a 2/5, it's hard to know who is right.
  • What it means: The doctors struggled to agree on how "clear" or "safe" the answers were. However, when they looked at the average of all five doctors, the difference between ChatGPT and Gemini on accuracy was still clear enough to see.

The "Unani" Twist

The study included questions about Unani medicine (a traditional healing system popular in South Asia) and cultural habits like praying on the floor. The researchers wanted to see if the robots understood these specific cultural contexts. While the study looked at these questions, the main conclusion was about the overall quality of the answers, not a specific ranking for how well they handled Unani medicine specifically.

The Bottom Line

If you are a patient with knee pain asking these robots for help:

  • ChatGPT tends to give you more complete and accurate information, like a detailed encyclopedia.
  • Gemini is slightly easier to read, but not by much.
  • Both are writing at a level that might be too complicated for the average person to understand easily.

The Final Warning: The researchers say that even though ChatGPT did better, neither robot should be used without a human doctor checking the work first. You shouldn't trust the robot's answer as the final word; a real doctor needs to review it to make sure it's safe and right for you.

In short: ChatGPT was the better "textbook," but both need a teacher (a doctor) to explain the lessons to the student (the patient).

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →