← Latest papers
💬 NLP

Measuring research data reuse in scholarly publications using generative artificial intelligence: Open Science Indicator development and preliminary results

PLOS and DataSeer developed a new LLM-based indicator that reveals a 43% research data reuse rate in scholarly publications, demonstrating that generative AI can measure open science impacts at scale and suggesting that the positive effects of data sharing are currently underestimated.

Original authors: Lauren Cadwallader, Iain Hrynaszkiewicz, parth sarin, Tim Vines

Published 2026-05-01
📖 4 min read☕ Coffee break read

Original authors: Lauren Cadwallader, Iain Hrynaszkiewicz, parth sarin, Tim Vines

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of scientific research as a massive, bustling kitchen. For years, we've been trying to figure out how often chefs (scientists) are actually using ingredients (data) left on the counter by other chefs, rather than just chopping up entirely new vegetables from their own gardens.

This paper is about a new, high-tech "smart camera" that PLOS (a scientific publisher) and DataSeer (a tech company) built to take a snapshot of this kitchen and count exactly how much ingredient-sharing is happening.

Here is the breakdown of their project in simple terms:

The Problem: The "Silent" Reuse

Previously, we tried to count how often scientists reused data in two main ways:

  1. Asking them directly: Surveys showed that about 50% of scientists say, "Yes, I use other people's data."
  2. Looking for formal citations: We looked for official "thank you" notes (citations) in papers. But this method was like trying to count guests at a party by only looking at people wearing nametags. It missed everyone who was just chatting in the corner. This method suggested reuse was very low (around 9–16%).

There was a big gap between what scientists said they did and what the "nametag" method could see.

The Solution: The AI "Super-Reader"

To solve this, the team built a Generative AI (a type of smart computer brain) that acts like a super-attentive reader. Instead of just looking for formal citations, this AI reads the entire text of a scientific paper to understand the story.

Think of it like a detective who doesn't just look for a specific fingerprint, but reads the whole diary to figure out if the person borrowed a book from the library or bought a new one.

How they trained the AI:

  • They showed the AI thousands of scientific articles and taught it to spot three things:
    1. Did the author create new data (like growing new vegetables)?
    2. Did the author reuse existing data (like using someone else's leftover ingredients)?
    3. Why did they do it? (The "reasoning").
  • They tested the AI against human experts to make sure it was accurate. The AI got very good at this, correctly identifying reuse about 85% of the time.

The Findings: The Kitchen is Busier Than We Thought

The AI analyzed 4,475 scientific papers published in the first three months of 2024. Here is what it found:

  • The Big Number: 43% of the papers reused data.
  • The Comparison: This is much higher than the old "formal citation" method (which saw only ~10%), but it lines up much better with what scientists told us in surveys (which said ~50%).
  • The Takeaway: The old methods were likely underestimating how helpful and collaborative scientists actually are. The "reuse" is happening, but it's often happening in the text of the paper without a formal "thank you" note attached.

Interesting Patterns

The AI also noticed some differences across the map and the subjects:

  • By Region: Scientists in Asia-Pacific were the most likely to reuse data (46%), while those in Africa were the least likely (31%). The authors suggest this might be related to how data is shared globally and concerns about who owns the data.
  • By Subject:
    • Physical Sciences and Social Sciences were the most balanced, reusing data almost as often as they created new data.
    • Life Sciences (like biology) were the most likely to create brand new data (76%) and the least likely to reuse existing data (37%).

Why This Matters

The authors argue that if we only look at formal citations, we are missing the full picture of how open science is working. By using this AI "smart camera," we can see that data sharing is having a bigger impact than we thought.

Important Note on Limits:
The authors are careful to say this is a "work in progress." They admit that:

  • It can be tricky to tell the difference between using your own old data versus someone else's.
  • They only looked at papers from one publisher (PLOS), so the results might look different in other journals.
  • The AI isn't perfect yet, but it's a huge step forward compared to the old ways of counting.

In short, this paper introduces a new tool that helps us see that scientists are sharing and reusing each other's work much more often than we realized, suggesting that the "open science" movement is working better than our old measuring sticks could show.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →