LLM assisted writing deserves empirical evaluation
This paper argues that LLM-assisted writing should be evaluated based on scholarly quality and accountability rather than treated as a detection problem, citing analysis of over 69,000 Health Informatics papers that links tool use to more focused presentation, broader citations, and globally distributed authorship.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Science has always been a conversation, one that relies on clear writing to share discoveries across borders and languages. For decades, researchers have depended on mentors, professional editors, and translation services to ensure their work is understood by the global community. Today, a new participant has joined this conversation: large language models. These are computer programs trained on vast amounts of text that can read, summarize, and rewrite human language with remarkable fluency. As these tools become common in laboratories and offices, a debate has emerged about their role in science. Many worry that using them might hide the true author of an idea, introduce errors, or lower the quality of research. Consequently, much of the current discussion focuses on how to detect when a machine has helped write a paper, treating the issue primarily as a matter of policing. However, this focus on detection may overlook a more important question: how does the actual use of these tools change the way science is communicated and who gets to participate in it?
To answer this, a team of researchers from Weill Cornell Medicine and Columbia University decided to look at the evidence rather than the theory. They examined a massive collection of 69,209 scientific papers in the field of health informatics, which deals with how data and technology are used in medicine and public health. The researchers split these papers into three groups to see how they differed. The first group contained papers written before late 2022, before the public release of the most advanced chatbots. The second group included papers written after that date that showed no signs of using these tools. The third group consisted of papers from the same recent period that the researchers identified as having been assisted by large language models. To find these assisted papers, they used a specialized computer program trained to recognize the subtle differences between human writing and text that had been rewritten or generated by an AI.
The analysis revealed that papers assisted by these tools were not the lower-quality, less serious work that critics often fear. Instead, the data suggested a different reality. These papers tended to be more focused in their presentation. While a typical scientific paper might wander through several different topics, the AI-assisted papers stuck more closely to a single, clear subject. Their titles and summaries were tightly aligned with the main theme of the research, making it easier for readers to understand exactly what the study was about and where it fit within the broader field. This clarity did not mean the ideas were less original; rather, it suggested that the tools helped authors position their work more precisely within established scientific conversations.
Another significant change appeared in how these papers referenced other work. The AI-assisted papers cited more sources than the papers written without assistance, and those sources covered a wider range of topics. This finding challenges the idea that machines simply make writing faster without adding value. It suggests that researchers are using these tools to explore literature they might not have otherwise found, perhaps by summarizing complex texts or generating new search terms. However, the researchers noted that having a longer list of references does not automatically mean the research is deeper or better. A bibliography can be broad but shallow, or it can include citations that are only loosely connected to the main argument. The tools make it easier to gather a wide net of information, but the responsibility for selecting the most relevant and accurate sources still rests with the human author.
Perhaps the most striking difference was found in who was writing these papers. The study showed that AI-assisted papers were more likely to have their first author based in a country where English is not the primary language. They also showed a higher level of collaboration between researchers on different continents and included more authors from lower- and middle-income economies. This pattern suggests that these tools may be acting as a bridge, helping scientists overcome the language barriers that often make publishing in international journals difficult and expensive. By providing on-demand help with editing and translation, the technology may be leveling the playing field, allowing more diverse voices to enter the global scientific conversation.
Despite these shifts in focus, citation habits, and authorship, the study found that the visibility and impact of the work remained steady. A common concern is that papers written with AI help might be ignored or cited less often. The data did not support this. When the researchers compared how quickly papers received citations after publication, the AI-assisted papers performed just as well as those written without help. They were also published in journals with high prestige, similar to the other groups. This indicates that the scientific community is not rejecting these papers; instead, they are being read and recognized at the same rate as traditional work.
The researchers concluded that the conversation about artificial intelligence in science needs to change. Rather than focusing on how to catch people using these tools or trying to separate "real" scholarship from "assisted" work, the scientific community should focus on the quality of the research itself. The study suggests that these tools are not inherently good or bad; they are simply a new part of the writing process, much like a spellchecker or a statistical software package. Their effect depends entirely on how they are used. If they help authors communicate more clearly and include more diverse perspectives, they can be a powerful force for good. If they are used to cut corners or hide errors, they can be harmful. The evidence shows that the tools are already reshaping how science is written and who writes it, and the best way forward is to judge the work by its accuracy and integrity, not by the software used to draft it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.