Evaluation of Large Language Models for Post-Cystectomy Sexual Health Counseling in Women: A Pilot Study
This pilot study evaluated three large language models for generating sexual health counseling after cystectomy, finding that ChatGPT produced the most guideline-concordant responses while all models generated content with excessive reading complexity, with adherence heavily influenced by prompt structure.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine you have three different "digital librarians" (ChatGPT, Gemini, and Perplexity) and you ask them to write a guide for women who are about to undergo a major bladder surgery called a cystectomy. Specifically, you want them to explain what will happen to their sexual health and relationships after the operation.
This study is like a report card for these librarians. The researchers wanted to see two things:
- Did they follow the rulebook? (Did they include all the important advice doctors are supposed to give?)
- Was the language too hard to understand? (Did they write like a professor or like a friendly neighbor?)
Here is what they found, broken down simply:
1. The "Rulebook" Test (Guideline Adherence)
The researchers gave the librarians a checklist based on official medical guidelines. They asked the librarians to answer six different questions about post-surgery life.
- The Winner: ChatGPT was the best student. It followed the rulebook the most closely, hitting about 77% of the required points.
- The Runners-up: Gemini and Perplexity scored much lower, hitting only about 50% and 46% of the points, respectively. They missed a lot of the key advice.
- The "Simple Question" Trick: The researchers noticed something interesting. When they asked the librarians using simple, everyday language (like a patient would ask), the answers were better and followed the rules more often. When they asked using complex, medical jargon (like a doctor would ask), the librarians actually got worse at following the rules. It's as if the librarians got confused by the fancy words and forgot the basics.
2. The "Reading Level" Test (Readability)
Even if the librarians got the facts right, it doesn't matter if the patient can't read the answer. The researchers checked how hard the text was to understand.
- The Problem: All three librarians wrote in language that was too hard. They wrote at a level suitable for college graduates or even people with advanced degrees.
- The Reality: Most people read at about a 6th-grade level. If you hand a 12th or 16th-grade reading level pamphlet to a patient who is already stressed about cancer surgery, they might not understand it at all.
- The Comparison: Perplexity wrote the most difficult, "graduate-level" prose. Gemini was slightly easier but still too complex. ChatGPT was the easiest to read of the three, but even it was still too hard for the average person.
3. The "One Big Idea" Discovery
The researchers used a special math tool (called Principal Component Analysis) to look at the different ways they measured "difficulty." They found that all the different tests for difficulty were actually measuring the exact same thing. It's like having five different rulers that all say the table is 6 feet long; they aren't five different measurements, they are just five ways of saying the same thing. This confirmed that the text was consistently too complex across the board.
The Bottom Line
Think of these AI tools as very knowledgeable but slightly clumsy assistants.
- ChatGPT is the most reliable assistant for remembering the important rules, but even it speaks in a language that is too fancy for a patient to easily digest.
- Gemini and Perplexity are less reliable with the rules and speak in even more complicated language.
- The Prompt Matters: If you ask these assistants in simple, plain English, they actually do a better job of following the rules than if you try to sound like a doctor.
The Takeaway: While these AI tools can be helpful, they currently produce information that is too complex for patients to understand and vary wildly in how much important medical advice they include. They are not ready to replace a doctor's conversation, especially for sensitive topics like sexual health after cancer surgery.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.