Large Language Models for Carotid-Cavernous Fistula Patient Education: A Multidimensional Comparative Study
This study evaluates and compares the quality, reliability, and readability of carotid-cavernous fistula patient education generated by ChatGPT, DeepSeek, and Doubao, finding that while these models offer useful information, they exhibit significant limitations in content completeness and linguistic complexity, necessitating professional oversight and further optimization.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a rare but serious condition where a high-pressure artery accidentally connects directly to a low-pressure vein behind the eye. This abnormal shortcut, known as a carotid-cavernous fistula, can cause the eye to bulge, turn red, and lose vision as blood backs up into the delicate tissues of the orbit. For patients facing this diagnosis, the path forward is often clouded by complex medical terms and a flood of scattered information online. They need clear, reliable answers about why this happened, how it is treated, and what to expect, but finding accurate guidance among the noise of the internet can be difficult. In recent years, people have turned to advanced computer programs that can hold conversations and generate text, hoping these tools might offer the clarity they seek. However, it remains unclear whether these artificial intelligence systems can explain such a specialized medical condition well enough to be trusted by patients or used to support doctors in teaching them.
A team of researchers set out to test exactly this question by asking three different artificial intelligence chatbots to answer twenty common questions that a patient with this eye condition might ask. The questions covered everything from the causes of the disease and the symptoms to watch for, to the details of diagnosis, treatment options, and daily life after care. The researchers fed these same questions into three popular language models, including two widely known international systems and one developed in China, ensuring that each bot answered without any extra help or internet browsing. They then gathered all sixty responses and had medical experts evaluate them blindly, meaning the doctors scoring the answers did not know which computer program wrote them. The experts judged the answers based on how complete and accurate the information was, how easy it was to read, and whether the advice was practical enough for a patient to use.
The study found that while all three computer programs could generate coherent and generally relevant information, they were not equally good at the job. One of the international models consistently provided higher-quality answers that were more reliable and better structured than the others. However, even the best responses had significant flaws. The most common problem was not that the computers made up facts, but that they left out crucial details. In many cases, the answers failed to explain the difference between types of the condition, did not clearly distinguish between minor symptoms and those requiring urgent care, or offered overly general advice that did not fit every patient's situation. The researchers noted that missing key information happened in about one out of every eight responses, a gap that could leave a patient confused about their own risk or the next steps they should take.
Another major finding concerned how difficult the text was to read. Medical information is already complex, but the researchers discovered that the language used by these computer programs was often too advanced for the average person. The answers frequently required a reading level equivalent to that of a college student, using long sentences and technical vocabulary that could overwhelm someone who is already stressed by a new diagnosis. One of the models produced responses that were significantly harder to read than the others, while even the best-performing model still used language that was more complex than what health experts typically recommend for patient education. This suggests that while the computers can access medical facts, they struggle to translate those facts into simple, everyday language that a worried patient can easily understand.
Despite these limitations, the study concluded that these tools are not useless; rather, they should be viewed as helpers rather than replacements for human doctors. The artificial intelligence systems showed they could provide a solid starting point for understanding the disease, but they cannot be trusted to give the final word on diagnosis or treatment plans. The researchers emphasized that for a condition as specialized as this, where the details of blood flow and eye pressure matter deeply, the information generated by computers must be checked and refined by a medical professional. The best path forward is to use these tools to supplement human care, ensuring that patients receive information that is not only accurate but also complete, clear, and tailored to their specific needs. Until the technology improves in its ability to simplify complex ideas without losing important details, the human doctor remains the essential guide in navigating the journey of this condition.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.