Verifiable by Construction: Claim-Level Evaluation of Verbatim Citation in Clinical Question Answering
This paper evaluates the ability of twelve large language models to generate verifiable clinical answers by attaching verbatim quotes to factual claims, revealing that while most models can successfully attach quotes to over 90% of claims, these quotes frequently fail to fully substantiate the details of the claims they accompany.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the high-stakes world of modern medicine, time is often the most scarce resource. When a doctor faces a complex patient case, they need answers that are not only accurate but also instantly trustworthy. For years, artificial intelligence has promised to help, offering to synthesize vast amounts of medical literature into clear, readable advice. However, a persistent problem has shadowed these tools: the tendency to invent facts, a flaw known as "hallucination." In a field where a single error can harm a patient, a doctor cannot simply trust a computer's word. They must be able to verify the source of every statement. Traditionally, when an AI provides an answer with a citation, that citation often points to a broad document, like an entire medical guideline or a research paper. This forces the busy clinician to hunt through hundreds of pages to find the specific sentence that supports the claim, a process that defeats the purpose of having an efficient assistant. The goal, then, is to build systems that do not just point to a source, but that prove their work by showing the exact words from that source right alongside the answer.
A team of researchers at Johns Hopkins University set out to test whether current artificial intelligence models could actually do this. They built a specialized system designed to answer clinical questions using four major medical guidelines covering heart disease, blood pressure, cholesterol, and diabetes. Instead of letting the models simply summarize information, the researchers instructed them to construct their answers sentence by sentence, attaching a direct, word-for-word quote from the original guideline to every single factual claim they made. To see if this worked, they created a rigorous testing framework that acted like a digital auditor. This system broke down every answer into its component claims and checked them against three specific standards: first, did the model attach a citation to the claim? Second, was the attached quote an exact, verbatim copy of the text in the original guideline, with no changes or paraphrasing? And third, did that specific quote actually contain enough information to prove the claim was true, without the reader needing to look anywhere else?
The researchers tested twelve different large language models, ranging from powerful, complex systems to smaller, faster ones, using 222 synthetic clinical questions derived from the guidelines. The results revealed a significant gap between what these models can do and what is required for true reliability. Most of the advanced models were surprisingly good at the first two steps. They successfully attached citations to over 90 percent of their claims, and in many cases, the quotes they provided were exact matches to the source text. However, the system faltered dramatically at the final, most critical step: proving the claim. While a model might provide a perfect quote, that quote often failed to support the full weight of the statement the model had made. For example, a model might quote a sentence saying a drug is "preferred" for a specific group, but then use that quote to support a broader claim that the drug is "required" for everyone. In one notable case, a top-tier model named Claude Opus 5 managed to attach a verbatim quote to 98 percent of its claims, yet only 37.1 percent of those claims were fully substantiated by the quotes alone. The quotes were there, and they were accurate, but they were insufficient to prove the point.
The study also highlighted that the size and type of the model mattered greatly. Smaller, lightweight models often failed to attach citations to their claims at all, or they provided quotes that were not exact copies of the source text, causing their verification rates to plummet. One of the smallest models tested, Claude Haiku, managed to fully support only about 25 percent of its claims. In contrast, the largest models performed much better, with the best systems supporting around 75 percent of their claims with exact, sufficient evidence. The researchers found that the primary reason for the failure was not that the models were lying or making things up, but rather that they were struggling to select the right evidence. They often grabbed a quote that was related to the topic but did not contain the specific details needed to back up the entire sentence. To test if the evidence was actually available, the researchers asked a separate AI to search the original documents for quotes that would have fully supported the claims. They found that for many models, the missing evidence was right there in the text they had already read; the models simply failed to extract and attach the correct pieces.
Ultimately, this work demonstrates that while artificial intelligence is becoming increasingly capable of retrieving and formatting medical information, it is not yet ready to be fully trusted without human oversight in a clinical setting. The ability to generate a fluent answer is distinct from the ability to build a verifiable one. The study shows that current models can often provide the raw materials for verification, but they frequently fail to assemble them into a complete, self-contained proof. The path forward requires systems that do not just cite sources, but that ensure every single word of their advice is backed by a specific, sufficient, and unaltered piece of evidence from the original guidelines. Until models can consistently bridge the gap between having a quote and proving a point, the responsibility for verification will remain firmly in the hands of the human clinician.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.