← Latest papers
📄 scientific communication and education

Who Funds Open Data Sharing? Analysis of data availability statements in biomedical publications

This study analyzes over 780,000 biomedical articles to reveal significant disparities in open data sharing compliance across funders and journals, finding an overall rate of 8.7% that rises to 24% for major funders and up to 92.9% for top journals, while highlighting that PDF-based detection methods identify 52% more sharing statements than XML-only approaches.

Original authors: Lawrimore, J., Li, C., Moraczewski, D., Poline, J.-B., Thomas, A.

Published 2026-07-20
📖 4 min read☕ Coffee break read

Original authors: Lawrimore, J., Li, C., Moraczewski, D., Poline, J.-B., Thomas, A.

Original paper dedicated to the public domain under CC0 1.0 (https://creativecommons.org/publicdomain/zero/1.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine the world of science as a massive, bustling library where researchers are constantly writing new books about how our bodies work, how diseases spread, and how to cure them. For a long time, these books were written in a way that kept the "secret ingredients" hidden. If a scientist wrote, "We found a cure using a special mix," they often didn't tell you what the mix was or how to make it. This made it hard for other scientists to check their work or build upon it. To fix this, a movement called "Open Science" started, championing a set of rules called FAIR principles. Think of FAIR as a promise that all the data behind a discovery should be Findable, Accessible, Interoperable (easy to mix with other data), and Reusable. The idea is simple: if you share your raw ingredients, everyone can taste the dish, check the recipe, and maybe even invent a better one. But here's the big question: Are scientists actually keeping their promise? Are they really sharing their data, or are they just saying they will?

This paper is like a giant, high-tech detective agency that went on a mission to find out exactly how many scientists are actually sharing their data. The authors, a team of data sleuths, didn't just ask researchers to raise their hands; they scanned nearly one million scientific articles published between January 2024 and June 2025. They used a clever two-part strategy to catch the truth. First, they looked at the digital "PDF" versions of the papers, which are like the final, polished books you'd see on a shelf. Second, they looked at the "XML" versions, which are like the raw, unformatted drafts used by computers. They found that the raw drafts often missed important notes about data sharing that were tucked away in the final polished versions. By using a smart computer program to read the PDFs, they could spot data-sharing promises that the other method missed.

So, what did they find? The news is a bit mixed. Overall, only about 8.7% of the articles they looked at actually had a clear statement saying, "Yes, our data is open and ready for you to use." That's a pretty low number, suggesting that for most scientists, the promise of open data is still just a promise. However, the story gets much more interesting when you look at who is funding the research and where it is being published. The authors found that the rate of sharing changes wildly depending on the "boss" behind the project.

Some big-money funders, like the French research agency or the Swiss National Science Foundation, saw sharing rates jump to around 22–24%. Even better, some specific journals (the places where papers are published) are leading the charge. Journals focused on genetics and brain science, like Nature Genetics, showed that up to 92.9% of their articles were sharing data. It's like a school where some teachers are strict about homework and the whole class gets A's, while other teachers are more relaxed and the class average is much lower. The paper suggests that when the rules are clear and enforced, scientists do share. But it also points out that we can't say for sure that the rules caused the sharing; maybe the scientists who like sharing just choose to work with those specific funders and journals in the first place.

The researchers also discovered that looking at the "draft" versions of papers (XML) was like trying to read a book with half the pages missing. Their PDF-based method found about 52% more data-sharing statements than the XML method alone. This means that for a long time, we might have been underestimating how much data is actually being shared, simply because we were looking at the wrong version of the text.

In the end, the paper doesn't declare a total victory for open science, nor does it say we've failed completely. Instead, it provides a clear, honest map of the landscape. It shows us that while the overall rate of sharing is still far from perfect, there are bright spots where it works incredibly well. The authors have built a public dashboard, like a live scoreboard, where anyone can check which funders and journals are doing the best job. Their work suggests that if we want to see more sharing, we need to look at what those top-performing groups are doing right and maybe try to copy their playbook. The journey to a fully open science library is still underway, but now we have a much better idea of where the doors are open and where they are still locked.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →