Evaluating Methods for Assessing LLM-Translated Clinical Discharge Instructions: A Comparison of Automated Metrics, MQM-Based Evaluation, and Clinician Review
This study compares automated metrics, MQM-based LLM evaluation, and bilingual clinician review for assessing LLM-translated clinical discharge instructions, finding that while these methods capture complementary aspects of quality, their low agreement underscores the continued necessity of targeted human review for ensuring accuracy in clinical meaning and terminology.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
When a patient leaves a hospital, the discharge instructions they receive are more than just a list of tasks; they are a critical bridge between professional care and daily life. If these instructions are not understood, the risk of medication errors, missed follow-up appointments, and worsening health conditions rises sharply. For patients who speak Spanish as their primary language, this bridge must be built in their native tongue. For decades, hospitals have relied on human translators to ensure these documents are accurate, but the rapid rise of artificial intelligence has offered a new, faster way to generate these translations. Large language models, the same type of technology that can write essays or answer complex questions, are now being tested to see if they can translate medical documents with the same care and precision as a human expert. The central question is not just whether the machine can translate words, but whether it can translate meaning, especially when a single wrong word could change a patient's understanding of their own health.
A team of researchers at the University of Colorado School of Medicine set out to answer this question by testing how well artificial intelligence translates hospital discharge instructions from English to Spanish. They did not simply ask the computer to translate; they built a rigorous experiment to see how different ways of asking the computer to work would change the results. They took two hundred real, de-identified discharge instructions from a large medical database and fed them into three different artificial intelligence models. To see if the way a human speaks to the machine matters, they used two distinct methods of prompting. In one method, they told the computer to act as an expert clinician in an intensive care unit, responsible for translating the text. In the other, they gave a much simpler, literal instruction to just translate the words, without any role or context. The goal was to see if framing the machine as a doctor would lead to better medical translations than treating it as a simple word-swapping tool.
To judge the quality of these translations, the researchers used three different approaches, each looking at the text from a different angle. First, they used automated computer metrics, which are standard tools in the field of language technology that count how many words and phrases in the translation match the original text. These tools are fast and can process thousands of documents, but they primarily look at surface-level similarities. Second, they used a structured framework called the Multidimensional Quality Metric, or MQM. This method breaks translation down into specific categories, such as whether the medical terms are correct, if the tone is appropriate for a patient, and if the grammar flows naturally. They had an artificial intelligence system act as a judge to score the translations across these categories. Finally, and most importantly, they brought in human experts. Two bilingual doctors who are fluent in both English and Spanish reviewed a random selection of the translations. They scored the documents using the same structured categories, serving as the gold standard for what a high-quality, safe translation should look like.
The results revealed a surprising disconnect between what the computers thought was good and what the human doctors thought was good. The automated metrics, which count word overlaps, consistently preferred the translations generated by the simple, literal prompts. When the computer was told to just translate word-for-word, the automated scores were higher, suggesting the output was closer to the original English text. However, when the researchers looked at the structured quality scores and the human reviews, a different picture emerged. The translations generated by the "clinician" prompt, where the computer was asked to act as a medical expert, often performed better in terms of style and tone, even if the automated word-counting tools rated them lower. This suggests that the tools used to measure translation success in the past might not be capturing the nuances that matter most in a hospital setting.
When the researchers compared the scores given by the artificial intelligence judges against the scores given by the human doctors, the agreement was surprisingly low. The computer and the human experts often disagreed on whether a translation was accurate or if the medical terminology was correct. The highest level of agreement between the machine and the human was found in areas like formatting and local conventions, where rules are clear and objective. However, in the most critical areas—such as the accuracy of the medical meaning and the appropriateness of the language for a patient—the machine and the human frequently saw things differently. Even the two human doctors did not always agree with each other, highlighting that judging medical translation involves a degree of human interpretation that is difficult to automate. The study found that while the artificial intelligence models produced generally high-quality translations, the areas where they struggled most were the very areas that require the deepest understanding of clinical context.
The researchers concluded that no single method is enough to ensure the safety of translated medical documents. The automated tools are useful for quickly checking large batches of text, and the structured scoring systems help organize the evaluation, but neither can fully replace the judgment of a human expert. The study suggests that the best approach is a layered one: using fast, automated checks to filter out obvious errors, followed by targeted reviews by human clinicians for the most critical parts of the text. This combination ensures that the efficiency of artificial intelligence is balanced with the necessary human oversight to protect patient safety. The findings serve as a reminder that in high-stakes environments like healthcare, the ability to translate words is not the same as the ability to communicate care, and that human expertise remains an essential part of the process.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.