Citation reliability of frontier large language models in medical writing and its automated verification
This study reveals that while frontier large language models generate a significant proportion of unreliable medical citations—primarily through misattribution rather than fabrication—an automated Chain-of-Verification system can detect these errors with expert-level accuracy, offering a viable solution for ensuring citation integrity in AI-assisted medical writing.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
In the modern practice of medicine, writing a research paper or a review article is a rigorous process that demands absolute precision. Authors must not only synthesize complex medical knowledge but also anchor every claim to a specific, existing source of evidence. These sources are citations, which act as the footprints of scientific discovery, pointing readers back to the original studies that support a new idea. For decades, this process has been a human endeavor, requiring careful checking of names, dates, and journal titles to ensure that a reference actually exists and says what the author claims it says. However, a new tool has entered the laboratory and the library: the large language model. These are powerful computer programs trained on vast amounts of text that can draft articles, summarize findings, and generate lists of references in seconds. While they offer a promise of speed and efficiency, a critical question has emerged: can these machines be trusted to find the right footprints, or do they sometimes invent them?
This question is not merely academic; it strikes at the heart of medical integrity. If a computer writes a medical review and includes a citation that looks real but points to the wrong study, or worse, to a study that never existed, the entire argument could collapse. This phenomenon, known as hallucination, has been a known weakness of earlier computer models. But as these tools have evolved, becoming faster and equipped with the ability to search the internet in real time, it has become unclear whether they have finally learned to cite correctly. The medical community needs to know if these new, advanced models are reliable enough to be used in serious research, and if they are not, whether there is a practical way to catch their mistakes before they are published.
A team of researchers set out to answer these questions by putting three of the most advanced computer models to a strict test. They asked each model to write a series of short medical reviews on nine different topics within the field of cardiology, the study of the heart and blood vessels. The topics ranged from common conditions to rare procedures, ensuring a mix of subjects with many published studies and subjects with very few. Each model was instructed to write a review of about 600 to 900 words and to include exactly thirty references for each article, for a total of 270 reviews and over 8,000 citations. Crucially, the researchers allowed the models to use their built-in web search tools, simulating how a user would actually employ them in a real-world setting. The goal was to see how many of these thousands of citations were actually correct.
The results revealed a landscape of significant variation and persistent error. When the researchers checked every single citation against the official medical database, they found that the models were not perfect. One model, GPT-5.5, performed the best, producing problematic citations in about 11.5 percent of cases. However, the other two models, Claude Opus 4.8 and Gemini 3.5 Flash, made errors in nearly 30 percent of their citations. This means that for every ten references generated by the less accurate models, roughly three were flawed. The researchers also investigated whether the models struggled more with topics that had fewer published studies, fearing that a lack of available information might confuse the computer. They found that while the error rates were slightly lower for topics with more literature, the difference was not strong enough to be considered a definitive rule. The models made mistakes across the board, regardless of how much information was available on the subject.
Perhaps the most revealing discovery was the nature of the mistakes themselves. The researchers expected to find many cases where the computer simply made up a fake article, a classic form of hallucination. Instead, they found that the vast majority of errors were far more subtle. About 77 percent of the problematic citations were cases of misattribution. In these instances, the computer provided a real, valid identifier for a real medical article, but that article was not the one the model claimed it was. It was as if the computer had found a real book in a library but placed it on the wrong shelf, labeling it with the title of a different book. This type of error is particularly dangerous because it looks correct at first glance; a human reader might see a valid number and assume the reference is trustworthy, only to discover later that the source does not support the point being made. True fabrication, where the computer invented an article that does not exist at all, was surprisingly rare, accounting for less than 1 percent of the errors.
Recognizing that these subtle errors are difficult for humans to catch without a systematic check, the researchers tested a new method for verification called Chain-of-Verification. This approach uses the computer model itself, but in a different way. Instead of asking the model to simply list references, the system breaks down the task into a series of small, independent questions. It asks the model to find the article based on its title and author, then to check if the journal and year match, and finally to confirm that the identifier points to the correct document. By forcing the model to verify each piece of information separately against the database, rather than relying on its memory, the system can catch its own mistakes. When the researchers tested this method against a set of references that had been carefully checked by a human expert, the automated system proved remarkably effective. It correctly identified 96.8 percent of the problematic citations, including every single case of misattribution and fabrication, while rarely flagging a correct citation as an error.
The study concludes that while the newest generation of medical writing tools is powerful, they are not yet ready to be trusted without supervision. The most common failure is not the invention of fake facts, but the mislabeling of real ones, a mistake that requires a specific kind of checking to detect. The researchers found that an automated verification process can perform this check with an accuracy comparable to that of a human expert. This suggests that the future of AI-assisted medical writing will not be a choice between human and machine, but a partnership where the machine drafts the content and a specialized verification system ensures that every single citation points to the truth. Until such verification becomes standard practice, the citations generated by these models should be treated with caution, as they may lead readers to the wrong source of truth.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.