Evaluating General-Purpose versus Medical-Specific Large Language Models for Perioperative Deep Vein Thrombosis Consultation: Perspectives from Physicians, Nurses, and Patients
This study demonstrates that general-purpose large language models outperform medical-specific models in generating high-quality, empathetic, and readable perioperative deep vein thrombosis consultation responses across evaluations by physicians, nurses, and patients, challenging the assumption that domain-specific labeling guarantees superior patient-facing information.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Every year, millions of people undergo surgery, and for many, the recovery period carries a hidden danger: deep vein thrombosis. This condition occurs when a blood clot forms in a deep vein, usually in the leg, often triggered by the immobility and physical stress of an operation. If the clot breaks loose, it can travel to the lungs and become life-threatening. Because the risk is so high, doctors and nurses spend considerable time teaching patients how to prevent these clots, what warning signs to watch for, and how to manage their care after leaving the hospital. Yet, in busy wards, finding the time to give every patient a personalized, clear, and reassuring explanation is a constant struggle. Traditional handouts are often too technical or dry, and verbal instructions can be forgotten or misunderstood.
In recent years, a new kind of digital assistant has emerged to help fill this gap. These are large language models, advanced computer programs trained on vast amounts of text that can hold a conversation and answer questions in human language. Some of these programs are built for general use, chatting about everything from history to cooking, while others are specifically tuned with medical knowledge to act as health consultants. The big question for hospitals and patients is simple: when it comes to explaining a serious medical condition like deep vein thrombosis, is a general conversationalist better, or does the specialist know more? A new study set out to find the answer by putting these digital tools to the test, not just for their accuracy, but for how well they connect with the people who need the information most.
The researchers gathered a team of fifteen people to act as judges: five doctors, five nurses, and five patients. They selected four different artificial intelligence programs to evaluate. Two of these were general-purpose models, the kind you might use for everyday questions, while the other two were medical-specific applications designed to give health advice. The team created nine standard questions that a patient might ask before or after surgery, covering topics like how to spot the early signs of a clot, how to use compression stockings correctly, and what to do if a dose of blood-thinning medication is missed. Each judge then asked these same nine questions to all four programs and rated the answers based on three criteria. First, they judged the overall quality and organization of the information. Second, they checked how accurate, clear, and practical the advice was. Third, and perhaps most importantly, they rated how empathetic the response felt—whether the computer sounded like it truly understood the patient's worry and fear.
The results surprised the researchers. In almost every category, the general-purpose models outperformed the medical-specific ones. The doctors, nurses, and patients all agreed that the general programs provided better-organized and more comprehensive answers. The difference was most striking when it came to empathy. The patients, in particular, felt that the general models sounded much more caring and human. They rated the general models significantly higher on empathy than the specialized medical apps. One of the medical-specific programs, for instance, produced answers that were so long and complex they required a high school education to understand, with sentences that dragged on for hundreds of words. In contrast, the general models kept their sentences shorter and their language simpler, making the information much easier to digest for an average person.
This finding challenges a common assumption that a tool built specifically for medicine will always be the best choice for medical questions. The study suggests that the medical-specific models may have been designed with such a strong focus on safety and caution that their answers became overly rigid, lacking the warmth and clarity that patients need during a stressful recovery. The general models, trained on a wider variety of human conversations, seemed better at balancing information with the emotional support that makes health education effective. Interestingly, the doctors, nurses, and patients did not disagree with each other on which models were better; they all saw the same pattern. The only exception was that the doctors were slightly less sensitive to the differences in empathy, likely because their medical training allowed them to focus more on the raw facts, while the patients and nurses felt the emotional tone of the answers more deeply.
Among the four programs tested, one general-purpose model stood out as the clear winner. It received the highest scores for quality, accuracy, and empathy from every group of judges. It managed to be both highly informative and easy to read, avoiding the trap of being either too simple or too complicated. The study does not claim that these artificial intelligence tools are perfect or that they can replace human doctors. Instead, it offers a clear guide for hospitals and nurses on how to choose the right digital assistant. The evidence suggests that for patient education, a general conversationalist that can speak clearly and kindly may be a more effective partner than a specialized medical bot that speaks in a language only a textbook could understand. As hospitals look to integrate these tools into daily care, the choice may not be about which label is more impressive, but about which tool actually helps the patient feel informed, safe, and understood.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.