← Latest papers
📄 medicine

Reliability and Clinical Accuracy of ChatGPT in Post-FESS Counselling: A Multi-Reviewer Evaluation.

This study evaluates ChatGPT's performance in generating postoperative FESS counseling, finding that while it offers acceptable reliability for standardized guidance with high inter-rater agreement on appropriateness, it exhibits limitations in contextual accuracy and readability that necessitate clinician oversight.

Original authors: PARAMITA DEBNATH, MISBAHUL HAQUE, PRAKRITI SAMADDAR, NEHA YADAV

Published 2026-07-05
📖 4 min read☕ Coffee break read

Original authors: PARAMITA DEBNATH, MISBAHUL HAQUE, PRAKRITI SAMADDAR, NEHA YADAV

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you just had a complex surgery to clear out your sinuses (called FESS). You leave the hospital with a million questions: "Is this bleeding normal?" "When can I go back to work?" "How do I clean my nose?"

In the past, you'd have to wait for a doctor to answer these. But now, people are turning to AI chatbots like ChatGPT for instant answers. This study asked a simple question: If a patient asks ChatGPT these specific post-surgery questions, is the robot giving safe, accurate, and easy-to-understand advice?

Here is the breakdown of what the researchers found, using some everyday analogies.

The Setup: The "Robot Doctor" Test

The researchers didn't just ask random questions. They created a strict "exam" with 15 specific questions that real patients usually ask after sinus surgery. These questions covered five main areas:

  1. Symptoms (e.g., "Is bleeding normal?")
  2. Nose Cleaning (e.g., "How do I wash my nose?")
  3. Lifestyle (e.g., "When can I go back to work?")
  4. Smell (e.g., "Will my sense of smell come back?")
  5. Medicine & Follow-up (e.g., "Do I need sprays forever?")

They fed these questions to ChatGPT (specifically the version available in early 2026) and then had four real ear, nose, and throat (ENT) doctors grade the robot's answers. The doctors gave scores from 1 (very poor) to 5 (excellent) on three things:

  • Appropriateness: Did the answer make sense for the situation?
  • Usefulness: Was it actually helpful to the patient?
  • Accuracy: Was the medical fact 100% correct?

The Results: The Robot's Report Card

1. The Overall Grade: B+ (Acceptable, but not perfect)
The robot got an average score of 3.9 out of 5.

  • The Good News: The robot was very good at being "appropriate" (4.1/5) and "useful" (4.0/5). It sounded like a helpful assistant and gave advice that felt right for a general situation.
  • The Bad News: Its "accuracy" was lower (3.6/5). While the advice was generally safe, it wasn't always medically precise enough to stand alone.

2. The "Textbook" vs. The "Real World"
The robot performed differently depending on the type of question:

  • The "Textbook" Questions (High Score): When asked about standard rules like "How do I clean my nose?" or "What medicine should I take?", the robot was excellent. These are like recipe instructions; they are written in books, and the robot can copy them perfectly.
  • The "Real World" Questions (Lower Score): When asked about lifestyle issues like "When can I go back to work?" or "How long until I feel better?", the robot stumbled. These questions are like asking a weather forecaster to predict if your specific picnic will be ruined. The robot doesn't know your specific job, your specific healing speed, or your specific health history. It has to give a generic answer, which can be vague or slightly off.

3. The "Language Barrier" Problem
The researchers checked how hard the robot's answers were to read.

  • The Problem: The robot wrote at a college reading level (Grade 12.8).
  • The Analogy: Imagine a teacher explaining a simple concept to a child, but using words like "utilize" instead of "use," and "subsequently" instead of "then." The robot's answers were too fancy. A typical patient might read the advice and nod along, but not truly understand it because the language was too complex.

4. The Doctors Agreed
When the four human doctors graded the robot, they mostly agreed with each other (especially on whether the advice was appropriate). This means the test was fair and the results are reliable.

The Bottom Line

The study concludes that ChatGPT is a good "supplementary" tool, like a study guide or a cheat sheet.

  • What it can do: It can reinforce standard instructions (like how to wash your nose) and remind you of general rules.
  • What it cannot do: It cannot replace the human doctor. It lacks the ability to look at you specifically and say, "Because your surgery was complex, you need to wait two weeks before working, not one."

The Takeaway: You can use the AI to get a quick, general overview, but you must still listen to your actual surgeon for the final, personalized verdict. The robot is a helpful assistant, but it is not the boss.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →