Rater choice determines measured fidelity: a multi-model evaluation of large language model raters for AI-generated plain-language biomedical summaries
This study demonstrates that the choice of large language model rater significantly influences measured fidelity and safety compliance in AI-generated biomedical summaries, with different models producing error rates that vary by a factor of four and leading to divergent conclusions about summary quality.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where the most important medical discoveries are locked behind a wall of complex language. Scientific papers are written for experts, filled with technical terms and dense data that the average person cannot decipher. Yet, for a patient trying to understand a new treatment for their condition, this barrier can be dangerous. If they cannot read the study, they cannot know the risks, the limitations, or the true benefits of a therapy. To bridge this gap, researchers and regulators increasingly rely on "plain-language summaries"—short, simple explanations of complex studies meant for the general public. As the volume of medical research explodes, writing these summaries by hand has become impossible. Instead, powerful computer programs known as large language models are being asked to do the work, translating dense science into readable prose in seconds. But a critical question remains: if a computer writes the summary, can another computer be trusted to check if it is accurate?
A recent study set out to answer this question by putting four different artificial intelligence models to the test. The researcher, who also built the tool that generated the summaries, created a controlled environment to see if these AI "judges" could reliably spot mistakes. The process began with fifty real, peer-reviewed medical articles covering five different diseases. A summarization application, designed to tailor content for different audiences like patients or professionals, turned these articles into 200 distinct plain-language sections. The goal was to see if the summaries were faithful to the original text. To do this, four different AI models—each a distinct type of large language model—were asked to grade the same 196 summaries using an identical, strict checklist. This checklist was anchored to the source text, meaning every question the AI had to answer was tied directly to a specific fact in the original article, such as "Did the summary mention this specific side effect?" or "Did it correctly state the study's limitations?"
The results revealed a startling inconsistency. When the four AI models graded the same summaries, they did not agree. In fact, their measurements of how many errors were present differed by a factor of four. Two of the models were very strict, finding a high number of errors in the text, while the other two were much more lenient, often finding almost no errors at all. The difference was so large that it mattered more than the type of disease being discussed or the intended audience for the summary. One model, in particular, found that more than half of the summaries contained no primary errors at all, while the strictest model found errors in nearly every single one. This divergence was not random noise; it pointed to a specific flaw in how the models operated. The lenient models seemed to be checking only what was present in the summary, missing the things that were supposed to be there but were missing. They failed to notice when a critical safety warning or a study limitation had been left out entirely.
This failure to detect missing information has serious implications for patient safety. The study focused heavily on "safety-critical omissions," such as warnings about who should not take a drug or what dangerous side effects might occur. In the source articles, twenty-eight sections contained at least one safety statement. The strictest AI model found that twenty-two of these summaries had omitted that safety information. The most lenient model, however, only spotted four of these dangerous omissions. When the researcher compared the AI results to a human expert who graded a smaller sample of the same summaries, the pattern became clear. The human expert found an error rate that sat between the strict and lenient AI models. The strict AI model closely matched the human's judgment, while the lenient models consistently missed more than half of the errors the human found. This suggests that relying on a lenient AI to check for safety could give a false sense of security, allowing dangerous summaries to pass through a quality gate undetected.
The study also uncovered a hidden technical problem that complicates the picture even further. During the testing, some of the AI models failed to return an answer at all, sending back empty responses. In the computer code used for the study, these empty responses were mistakenly treated as perfect scores, as if the model had read the text and found it flawless. When the researcher re-ran these specific tests, they discovered that many of the "perfect" scores were actually just silent failures where the model had given up. Even when accounting for these failures, the conclusion remained the same: the lenient models were systematically under-detecting errors. The study concludes that there is no such thing as a universally reliable AI judge. A model that works well for one task or one type of text may fail completely at another. The only way to know if an AI is trustworthy for checking medical summaries is to test it against a human standard on the specific type of content it will be grading. Without this validation, the choice of which AI to use as a judge can change the entire conclusion about whether a summary is safe to read.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.