Performance of a domain-specific large language model in answering patient questions in psychiatry
Although the domain-specific LLM "MIND" outperformed general chatbots in objective rubric scores and completeness, board-licensed psychiatrists ultimately preferred the responses generated by ChatGPT despite MIND's superior clinical fidelity training.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern landscape of healthcare, patients increasingly turn to the internet for answers about their medications, seeking clarity on how a drug works, what side effects to expect, or how it might interact with their daily lives. While large language models—powerful computer programs capable of generating human-like text—have shown promise in answering medical questions, they often struggle with the specific demands of clinical care. These general-purpose tools can produce fluent and informative answers, but they sometimes lack necessary nuance, omit critical safety warnings, or invent facts that could mislead a vulnerable person. In the field of psychiatry, where public information is already inconsistent and medication adherence is a major challenge, the risk of receiving inaccurate advice is particularly high. Doctors often have only fifteen to twenty minutes per visit, leaving little time to address every patient question, which drives many to seek information from unverified online sources. The central question for researchers has been whether a specialized computer model, trained only on trusted medical sources, could provide the accurate, safe, and personalized guidance that patients need without the errors common in general chatbots.
To investigate this, a team of researchers at Lurie Children's Hospital of Chicago developed a specialized artificial intelligence system they named MIND. Unlike general chatbots that are trained on vast, uncurated portions of the internet, MIND was built using a much narrower and safer approach. The researchers fed the system exclusively into a collection of patient education materials from seven authoritative sources, including the American Psychiatric Association, the Food and Drug Administration, and the National Library of Medicine. The system was designed with strict safety principles: it was programmed to retrieve answers only from these verified documents, to cite its sources, and to clearly state when a patient should consult a doctor rather than relying on the computer. The researchers tested this new tool against two other well-known systems: a widely used general-purpose chatbot and a medical decision-making platform designed for doctors but not the public. They asked all three systems fifty common questions about a specific antidepressant medication called escitalopram, covering topics from how the drug works and how to take it, to its side effects and safety for special groups like pregnant women or the elderly.
The results of the comparison revealed a distinct advantage for the specialized model in terms of reliability and completeness. When the responses were evaluated by a computerized rubric that checked for accuracy, clarity, and safety, the MIND system outperformed both the general chatbot and the doctor-focused platform in every category. It provided more complete answers, acknowledged uncertainty more effectively, and was better at knowing when to refer a patient to a physician. However, the story became more nuanced when the researchers asked ten board-certified psychiatrists to review the answers. While the doctors agreed that the specialized model was safer and more complete, they found that the general chatbot was slightly more accurate in its specific claims. Perhaps most surprisingly, the psychiatrists overwhelmingly preferred the responses from the general chatbot over the specialized one. This preference was especially strong among younger, early-career doctors. The researchers noted that the general chatbot tended to give shorter, punchier answers that felt more direct, whereas the specialized model provided longer, more detailed responses that sometimes included information the doctors felt was unnecessary or slightly inaccurate in the context of a brief summary.
The study highlights a complex trade-off in the design of medical artificial intelligence. The specialized model succeeded in its primary goal of being grounded in verified facts and avoiding the hallucinations common in general systems, yet it struggled to match the brevity and style that human experts found most appealing. The researchers observed that as the answers became more detailed and complete, they were sometimes rated as less accurate, suggesting that adding necessary context can make an answer feel less precise to a human reader. Furthermore, the specialized model could not answer every question; for twelve of the fifty queries, it correctly admitted it did not have the information, whereas the other systems attempted to answer all of them. This refusal to guess was a deliberate safety feature, but it also meant the tool was less versatile than its competitors. The psychiatrists also noted that the specialized model's responses were written at a higher reading level than the general chatbot's, a reflection of the fact that the source documents themselves were written at a complex level, far above the recommended reading grade for patient materials.
Ultimately, the research suggests that while a specialized, restricted-corpus model can provide a safer and more complete foundation for patient education than a general chatbot, it is not yet a perfect replacement for human judgment or the preferred communication style of clinicians. The specialized system demonstrated that it is possible to build an artificial intelligence that adheres strictly to medical evidence and avoids dangerous errors, but it also revealed that achieving high accuracy and completeness can sometimes come at the cost of readability and user preference. The authors conclude that their work serves as a template for building future tools that can enhance psychiatric care, but they emphasize that these systems require further optimization to balance safety, accuracy, and the way humans actually prefer to receive information. The path forward involves refining these models to be as reliable as the specialized system but as clear and engaging as the general chatbots, ensuring that patients receive the high-quality, evidence-based guidance they need without confusion or delay.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.