What does Retraction Watch's "computer-generated content" reason capture? A metadata analysis of 9,161 retractions
This metadata analysis of 9,161 retractions reveals that the "computer-generated content" reason in Retraction Watch is heavily concentrated in a single publisher and strongly associated with paper mills, but these patterns reflect co-coding practices rather than providing definitive evidence about the actual presence of generative AI in the retracted articles.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the vast library of human knowledge, every published study is supposed to be a building block of truth. But sometimes, blocks are pulled out. When a journal decides a paper is flawed, fraudulent, or simply wrong, it issues a retraction, effectively pulling the article from the shelves of science. For decades, researchers have tracked these retractions to understand how the system fails. Recently, a new reason has appeared on the retraction notices: "computer-generated content." This label is applied when an investigation finds that a paper was written, in whole or in part, by a machine rather than a human. As artificial intelligence tools have become more powerful and accessible, this reason has grown rapidly, leading many to believe it signals a new wave of misconduct driven by generative AI. The question is whether this label truly points to a surge in AI-written papers, or if it is capturing something else entirely.
A researcher named Jonas Mandalunes set out to investigate exactly what this label captures. Using a massive database of over 66,000 retraction records, the study focused on the 9,161 papers that carried the "computer-aided or computer-generated content" reason. The goal was to see if these papers behaved differently from other retracted articles in terms of how long they stayed in print before being pulled, what other problems they had, and how often they were cited after their retraction. The findings reveal that the story behind this label is far more complex and specific than a simple surge in AI misuse.
The first thing the data showed was a massive spike in numbers, but it was not a steady climb across the entire world of science. Before 2020, this reason was almost non-existent, appearing on only 28 papers. By 2023, however, it accounted for nearly 42 percent of all retractions recorded that year. This sudden explosion looked like a tidal wave of AI-generated fraud. Yet, when the researcher looked closer, the wave was not coming from everywhere. It was concentrated almost entirely in one specific place: a single publisher known as Hindawi. In fact, 70 percent of all the papers with this label came from that one publisher. This publisher had recently undergone a massive internal investigation into its special-issue programs, uncovering widespread manipulation of the peer-review process. The study suggests that the label is largely capturing the aftermath of this specific, large-scale cleanup rather than a general epidemic of AI writing across all scientific fields.
Another common assumption was that these computer-generated papers were being caught much faster than other fraudulent work. Because machines can churn out text quickly, it seemed logical that their mistakes would be spotted quickly too. The raw numbers seemed to support this, showing that the "computer-generated" papers had a much tighter range of time before retraction compared to the messy, varied timelines of other retracted papers. However, this tight timing was an illusion created by the concentration of the data. When the researcher compared these papers only to other papers from the same publisher and the same years, the difference in timing vanished. The "computer-generated" papers were not being caught faster; they were just part of a batch of papers that were all reviewed and retracted at the same time during that specific publisher's investigation. Once the dominant publisher was removed from the comparison, the idea that these papers have a unique, fast-tracked timeline disappeared completely.
The study also examined the connection between this label and "paper mills," which are organizations that sell fake authorship and fabricated manuscripts. Initially, the data showed a strong link: 71 percent of the "computer-generated" papers also carried a paper-mill flag, compared to only 9 percent of other retracted papers. This suggested that the label was a reliable marker for these criminal operations. But again, the picture changed when the data was sorted by publisher and year. Before 2023, the link was non-existent. In 2023 and 2024, the link became very strong, but only within specific publishers. In other large publishers, the relationship actually flipped, with the "computer-generated" papers being less likely to be linked to paper mills than other retracted papers. This extreme variation means there is no single, universal rule connecting the label to paper mills; the association depends entirely on who is doing the investigating and when.
Finally, the researcher looked at whether these papers were cited less often after being retracted, which would indicate that the scientific community was paying attention to the warning. At first glance, it appeared that the "computer-generated" papers were cited significantly less than similar retracted papers. But this difference was largely due to the fact that the control group of papers had been cited more heavily before they were even retracted. When the researcher matched the papers perfectly based on how many citations they had already received, the difference in post-retraction citations almost disappeared. The remaining tiny gap was not statistically significant, suggesting that the scientific community treats these papers no differently than other retracted work once their prior reputation is accounted for.
Perhaps the most telling finding came from a search for specific words in the retraction notes. The researcher looked for mentions of specific AI tools, like "ChatGPT" or "large language models," in the metadata of the papers. Out of thousands of records, only 21 contained such specific terms. Even more surprisingly, none of the papers that carried the "computer-generated content" label and also mentioned a specific AI tool were actually about the tool itself; they were studies about AI, not papers written by AI. The one paper that did mention AI-generated figures did not even carry the "computer-generated content" label. This indicates that the label is not being applied because the retraction notices explicitly say "this was written by a machine." Instead, the label is being applied based on broader investigations into the integrity of the paper, often alongside findings of compromised peer review.
The study concludes that the "computer-generated content" label in the retraction database is not a reliable proxy for generative AI misuse. It is a record of a specific, massive investigation into a single publisher's special-issue programs, where the label was applied to a wide range of issues including paper mills and peer-review failures. The data shows that the label does not point to a unique type of paper that is caught faster or treated differently by the scientific community. Until researchers can read the actual retraction notices and verify the content of the papers, using this label as a measure of AI-generated fraud will lead to a distorted view of the problem. The real story is not a sudden invasion of AI-written science, but a concentrated cleanup of a specific publishing workflow that happened to use a broad label to describe its findings.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.