Validation of Large Language Models in Healthcare: a Scoping Review
This scoping review maps existing validation methods for Large Language Models in healthcare, categorizing them by task and evaluation type to provide stakeholders with a structured framework for selecting context-appropriate strategies before clinical implementation.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern world of medicine, a new kind of tool has emerged that can read vast libraries of medical records, answer complex questions about symptoms, and draft patient letters with a speed no human can match. These tools, known as large language models, are a form of artificial intelligence trained on enormous amounts of text. They do not simply retrieve facts from a database; they generate new sentences, weaving together information to create responses that sound natural and human. While this technology promises to relieve doctors of administrative burdens and help patients understand their care, it carries significant risks. Because these models create text rather than just recalling it, they can sometimes invent facts that sound plausible but are entirely false, a phenomenon often called hallucination. They might also reflect biases present in the data they were trained on, potentially leading to unfair or harmful advice. Before such a powerful tool can be trusted in a hospital or a clinic, it must be rigorously tested to ensure it is safe, accurate, and fair.
A team of researchers at the University Medical Center Utrecht in the Netherlands set out to map the landscape of how these models are currently being tested. They conducted a comprehensive review of scientific literature to find every method used to validate large language models in healthcare. The team looked at studies published from 2005 up to late 2024, searching for papers that described how to check the quality of these AI systems. They focused on five specific tasks where these models are most likely to be used: general text generation, summarizing long documents, answering questions, extracting specific medical details from text, and acting as conversational agents like chatbots. After screening thousands of potential studies, they identified 266 relevant papers that offered concrete methods for evaluation.
The researchers organized these validation methods into three main approaches. The first is human evaluation, where people read the AI's output and judge its quality based on criteria like correctness, relevance, or tone. This is often considered the gold standard because humans understand nuance and context better than machines. The second approach is automatic evaluation with a human reference, where a computer program compares the AI's answer to a perfect answer written by a human, measuring how closely they match. The third is fully automatic evaluation, where the computer judges the output on its own, without any human-written example to compare it against. The review found that while many methods exist, they are often scattered and specific to certain tasks. For instance, checking if a summary is good requires different tools than checking if a chatbot is being polite.
One of the most important findings is that simple, surface-level checks are not enough. The review highlights that methods which merely count how many words overlap between the AI's answer and a reference text are often misleading. These traditional metrics fail to capture whether the meaning is correct or if the information is safe for a patient. Instead, the researchers found a growing shift toward more sophisticated techniques. These include methods that check if the AI's statements are factually consistent with the source material, ensuring it does not invent medical facts. There are also tools designed to test if the AI behaves consistently when the same question is asked in slightly different ways, or if it treats people from different backgrounds fairly. The review emphasizes that for high-stakes fields like healthcare, a single score is rarely sufficient. A robust validation strategy usually requires a combination of methods, checking for accuracy, safety, fairness, and clarity all at once.
The authors also point out a significant gap in current practices. While many studies propose new ways to test these models, there is a lack of standardized guidelines specifically for healthcare. Many existing methods were developed for general language tasks and may not catch the specific dangers of medical errors. The review suggests that relying solely on automated tools is risky because they can sometimes miss subtle errors that a human would catch. Conversely, relying only on human judges is slow and expensive. The most promising path forward appears to be a hybrid approach, where automated tools handle the bulk of the checking, but human experts step in to verify the results and ensure the model aligns with medical standards.
Ultimately, this review serves as a roadmap for anyone looking to use these powerful tools in medicine. It does not claim to have solved the problem of validation, but rather provides a clear inventory of the tools currently available. The researchers conclude that as these models become more common, the medical community needs to move beyond generic tests and adopt structured, task-specific validation frameworks. By understanding the strengths and weaknesses of each evaluation method, developers and clinicians can better ensure that these artificial intelligence systems are reliable partners in patient care, rather than unpredictable sources of error. The work underscores that validation is not a one-time event but an ongoing process essential for the safe integration of technology into the delicate ecosystem of healthcare.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.