Detecting Evidence of LLM Contributions to Open Access Biomedical Literature: A Bibliometric Analysis
This bibliometric analysis of nearly 3 million open-access biomedical articles reveals a significant post-2022 increase in LLM-associated terminology, indicating rapid adoption of generative AI tools and highlighting the urgent need for transparency and detection methods to preserve scientific integrity.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
For decades, scholars have relied on tools to make their writing clearer and their research more efficient. Spell-checkers and grammar software have long been standard helpers, smoothing out errors and suggesting better word choices without changing the core voice of the author. In recent years, a new generation of these tools has emerged, capable of generating entire paragraphs of text that sound remarkably human. These systems, known as large language models, are trained on vast amounts of written material and can produce coherent, polished prose on almost any topic. While they offer the promise of speeding up the writing process, they also raise a difficult question for the scientific community: how can we tell the difference between a sentence written by a human researcher and one crafted by a machine? This distinction matters deeply because the integrity of scientific literature depends on the authenticity of the ideas and the accuracy of the facts presented within it. If a significant portion of new research is being drafted or heavily edited by these tools, the way we read, trust, and build upon that knowledge may need to change.
A team of researchers set out to answer this question by looking directly at the words themselves. They focused on the vast collection of open-access medical and biological research papers available in a public database called PubMed Central. This collection contains millions of articles, providing a massive window into how scientists across the globe are writing their work. The team did not try to guess which specific papers were written by machines, a task that has proven difficult for both humans and computer programs. Instead, they looked for a different kind of signal. They examined whether certain words and phrases, which previous studies had found to be used far more often by artificial intelligence than by people, had suddenly become much more common in the scientific literature after late 2022, when these tools became widely available to the public.
The researchers gathered nearly three million articles published between 2019 and 2024. They created a list of 190 specific words and phrases that had been flagged in earlier research as being characteristic of machine-generated text. These included adjectives like "meticulous" and "intricate," verbs like "delve" and "underscore," and longer phrases such as "navigating the complexities of" or "it is important to note." They then scanned every single article in their dataset to see how often these terms appeared. The goal was to see if the frequency of these words changed over time, specifically looking for a sharp jump in usage after the release of the new AI tools.
The results showed a clear and dramatic shift in the language of biomedical research. Before 2022, the use of these specific terms grew slowly and steadily, consistent with normal changes in how people write. However, starting in 2023 and continuing into 2024, the usage of many of these words exploded. For instance, the word "meticulously," which appeared less than a thousand times in the entire database in 2019, showed up more than 16,000 times in 2024. The phrase "it's important to note" saw a similar surge, jumping from just 29 instances in 2019 to nearly 800 in 2024. Perhaps most telling were the longer, more complex phrases. The combination "delving into the intricacies of" was virtually non-existent in the literature before 2023, yet by 2024 it had appeared in dozens of articles. The researchers noted that these were not just isolated words becoming popular; the specific combinations of words that are known to be hallmarks of AI writing were appearing together with increasing frequency.
The study also looked at how these words were being used in combination with one another. They found that pairs and groups of these AI-associated terms were appearing together in the same articles at rates that were statistically unlikely to happen by chance. For example, the words "artificial intelligence" and "underscore" began appearing together in the same papers far more often than before. The researchers observed that these patterns were not just a slow evolution of language style, which would happen gradually over many years, but rather a sudden inflection point that coincided exactly with the public release of generative AI tools. While the study cannot prove that any single article was written by a machine, the sheer scale of this linguistic shift suggests that these tools are being used extensively by authors across the field.
The authors of the study emphasize that this does not mean the research itself is necessarily flawed, nor does it suggest that every paper using these words is inauthentic. Instead, the findings point to a rapid and widespread adoption of new writing aids that are changing the texture of scientific communication. The researchers argue that because these tools can sometimes make mistakes or produce content that lacks true understanding, the scientific community needs to develop a better way to identify when they are being used. They suggest that rather than trying to ban these tools, which are likely here to stay, scholars, editors, and reviewers should focus on transparency. By acknowledging when and how these tools are used, the scientific record can remain trustworthy even as the methods of creating it continue to evolve. The study concludes that the language of science is changing, and recognizing these new patterns is the first step toward understanding how artificial intelligence is reshaping the way we share knowledge.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.