Large Language Model Responses to Patient Questions About Return to Activities of Daily Living After Knee Arthroplasty: A Multidimensional Evaluation Before and After Plain-Language Prompting
This study evaluates how plain-language prompting affects Turkish responses from ChatGPT, Gemini, and Claude regarding return to daily activities after knee arthroplasty, finding that while readability and understandability improved, completeness and treatment-information quality often declined, highlighting the need for professional review before using such AI-generated guidance.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
After a major knee replacement surgery, the path to recovery is not just about healing the joint; it is about reclaiming a life. Patients face a long list of practical questions: When can they drive? How soon can they climb stairs? What daily chores are safe to attempt? While doctors provide the initial roadmap, the time spent in the hospital is often too short to answer every specific worry that arises once a person returns home. Consequently, many turn to the internet for guidance, increasingly asking artificial intelligence systems for help. These computer programs, known as large language models, are designed to understand human language and generate answers that sound like a knowledgeable conversation partner. They promise to fill the gaps in patient education, offering immediate, accessible advice on when and how to return to normal activities.
However, a new study from researchers in Turkey asks a critical question: Is the advice these machines give actually good enough to trust? The researchers wanted to know if these artificial intelligences could provide safe, complete, and clear instructions for people recovering from knee surgery, and whether asking them to "speak simply" would make their answers better or worse. They focused on three popular systems—ChatGPT, Gemini, and Claude—and tested them using real questions that patients in Turkey were actually typing into search engines. The goal was to see if these tools could act as reliable partners in recovery or if they might leave patients with incomplete or confusing information.
The researchers set up a careful experiment to test these systems. They gathered twelve common questions that patients ask after knee surgery, covering topics like managing pain, getting dressed, and moving around the house. They asked each of the three artificial intelligence models these questions in their original form. Then, they asked the same models to answer the questions again, but this time they added a specific instruction: "Please explain this in a way that is easier to understand." This second step was designed to see if forcing the machines to use plain language would improve their helpfulness. Two physical therapists, who did not know which model produced which answer, then read every response. They scored the answers based on several criteria: whether the medical facts were correct, whether the advice was safe, how complete the information was, and how easy it was to read and act upon.
The results revealed a complex picture of strengths and weaknesses. On the most vital points, the artificial intelligence systems performed very well. Every single response they generated was medically accurate and safe, meaning none of the machines gave advice that would harm a patient or contradict established medical knowledge. This is a significant finding, suggesting that for general safety, these tools are already quite reliable. However, when the researchers looked at how complete the answers were, the story changed. When the models were asked to simplify their language, two of them—ChatGPT and Claude—actually became less complete. They dropped important details, such as specific timelines for recovery or necessary precautions, in an effort to sound simpler. The third model, Gemini, managed to keep its answers complete even after simplifying the language, performing better than the others in this specific area.
The study also found that while the language became easier to read, the quality of the information regarding treatment choices suffered. The scores for how well the answers explained different treatment options and supported decision-making dropped significantly for ChatGPT and Claude after the simplification request. It appears that when these machines try to speak more simply, they sometimes strip away the nuance and depth that patients need to make informed choices. Furthermore, while the answers became easier to read, they did not necessarily become more actionable. The researchers found that the instructions remained only about sixty percent actionable across all models. This means that while the text was clear, it often lacked the specific, step-by-step details a patient needs to actually perform a task safely, such as exactly how to move a limb or how often to perform an exercise.
The researchers concluded that while these artificial intelligence tools can provide accurate and safe general information, they are not yet ready to replace the personalized advice of a physical therapist. Asking a machine to speak in plain language makes the text easier to read, but it can also cause it to lose the critical details that make the advice useful and complete. The study suggests that these tools can serve as a helpful supplement to patient education, but their output must be reviewed and tailored by a healthcare professional before a patient uses it to manage their own recovery. The best approach remains a partnership where technology provides the initial information, and a human expert ensures it fits the specific needs and safety requirements of the individual patient.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.