← Latest papers
💻 computer science

A million-scale cross-document clinical citation-evidence question answering dataset

This paper introduces RefQA, a million-scale dataset of 1.52 million cross-document clinical citation-evidence question-answer records derived from PubMed Central, featuring structured metadata, verifiable claim grounding, and high-quality annotations to support research in clinical citation analysis.

Original authors: Imjin Ahn, Young-Hak Kim, Sanghyun Park, Tae Joon Jun

Published 2026-09-23
📖 5 min read🧠 Deep dive

Original authors: Imjin Ahn, Young-Hak Kim, Sanghyun Park, Tae Joon Jun

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the vast library of medical science, where millions of research papers are published every year, a single sentence often carries the weight of a life-saving decision. When a doctor reads a new guideline recommending a specific treatment, that recommendation usually rests on a chain of references pointing back to earlier studies. These references are the links in the chain, connecting a current claim to the original evidence. However, for decades, the scientific community has known that these links are not always strong. Sometimes a paper claims to support an idea, but the original study it cites actually says something different, or perhaps nothing about that topic at all. This gap between what a paper says it found and what the source actually contains creates a dangerous uncertainty for clinicians trying to make sense of the literature. To fix this, researchers needed a way to read not just the papers, but the connections between them, verifying that every claim is truly backed by the evidence it cites.

A team of researchers has now built a massive digital tool to solve this problem, creating a dataset containing over 1.5 million records of these connections. They call it RefQA. The project focuses on the clinical literature, specifically the papers that deal with human health and patient care. The team started by gathering millions of full-text articles from a public archive known as PubMed Central. They did not just look at the text; they looked at the specific moments where one paper mentions another. In the digital files of these articles, every time a writer cites a previous study, there is a hidden marker that links the two documents. The researchers used computer programs to find every single one of these links, creating a map of who cites whom across the entire field of clinical medicine.

The challenge was to understand the meaning behind each link. A citation is not just a pointer; it is a statement of belief. When a writer says, "As shown in study X," they are making a claim about what study X proved. The researchers wanted to know if that claim was true. To find out, they used a powerful artificial intelligence system to read the citing sentence and the abstract of the cited paper side by side. The AI was trained to act like a careful editor, asking specific questions: What is the writer trying to say? Does the original paper actually support that idea? Is the evidence strong, weak, or missing? The system was instructed to be extremely precise. If the AI decided the original paper supported the claim, it had to pull a direct, word-for-word quote from the original abstract to prove it. This requirement meant that the truth of the claim could be checked instantly by anyone reading the record.

The team applied a strict filter to ensure the data was relevant to human medicine. They kept only the citations where both the writing paper and the cited paper were about clinical research involving patients. This rule removed millions of records that mixed basic science experiments with human studies, ensuring that the final dataset contained only chains of evidence that doctors could actually use. After running this process on over 1.5 million citation events, the result was a clean, organized collection of records. Each record contains the context of the citation, the details of the two papers involved, and the AI's analysis of the relationship between them. The system was remarkably accurate, with nearly every record following the required format correctly. When human experts checked a small sample of the work, they agreed with the AI's judgments about whether a citation was relevant or misleading in most cases, confirming that the machine had learned to spot the difference between a solid link and a broken one.

One of the most important features of this dataset is its ability to catch errors. The researchers found that a significant number of citations in the medical literature do not actually support the claims made about them. In some cases, a paper might cite a study to back up a statistic, but the study never mentioned that number. In others, a paper might claim a treatment works, while the cited study actually found no difference between the treatment and a placebo. The dataset captures these mismatches, labeling them clearly so that future computer programs can learn to identify them. The researchers also provided a version of the data that is free for anyone to use in commercial projects, while keeping the full collection available for those who need to see the entire picture, including papers with more restrictive usage rules.

The scale of this work is unprecedented. Previous attempts to map these connections were either too small to be useful for training advanced computer systems or too broad to capture the specific meaning of the citations. This new dataset bridges that gap, offering a million-scale resource that is both deep and wide. It allows researchers to train new artificial intelligence models to read medical literature with a level of scrutiny that was previously impossible. By teaching these models to verify claims against their sources, the work helps build a future where medical decisions are based on evidence that has been rigorously checked. The dataset is now available for scientists and developers to use, providing a foundation for tools that can help doctors navigate the overwhelming volume of medical research with greater confidence and accuracy.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →