When AI Writes, Who Gets Cited? Evidence of Citation Monoculture Across Language Models
This study demonstrates that diverse language models exhibit a "citation monoculture" by disproportionately selecting the same narrow subset of real papers due to shared content-level preference filters, a systemic bias that persists even when hallucinations are eliminated and cannot be resolved merely by equalizing retrieval or mixing vendors.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
For decades, the way science advances has relied on a simple, human habit: researchers read the work of others and decide which papers to mention in their own writing. This act of citation is more than just a formality; it is the currency of scientific credit. When a paper is cited often, it gains visibility, leading to more citations in a cycle known as the Matthew effect, where the rich get richer. This dynamic has always depended on human readers, who can see an author's reputation, a journal's prestige, or a paper's past success, and let those signals guide their choices. But a new force is entering the room. Artificial intelligence is moving beyond simply checking grammar to drafting entire literature reviews and selecting the references that shape what the scientific community sees. The question is no longer just whether these machines can write, but what they choose to read. If an AI is asked to find the most important work on a topic, does it simply mimic human bias, or does it create a new kind of bias all its own?
A team of researchers set out to answer this by building a controlled experiment that stripped away everything that usually influences a citation. They gathered 120 real scientific papers on a specific topic, but before showing them to the AI, they hid all the usual status signals. The authors' names were replaced with fake ones, the publication years were shuffled, and the number of times each paper had been cited was completely removed. The AI models were then asked to write short reviews or position papers, but with a strict rule: they could only mention up to ten of the thirty papers shown to them in each round. This scarcity forced the models to make real choices rather than listing everything they saw. To ensure the results were not just a fluke, the researchers compared the AI's choices against a baseline where a computer picked papers at random from the same list, and they also asked eight human experts to make the same choices under the exact same conditions.
The results revealed a striking and uniform pattern. While the human experts distributed their citations fairly evenly, treating the papers much like a random picker would, every single AI model tested concentrated its attention on a tiny, narrow subset of the papers. In the human group, the top ten percent of papers received about 19 percent of the citations, which is close to what you would expect by chance. In contrast, the AI models gave between 23 and 30 percent of their citations to just the top ten percent of papers. More importantly, the models consistently ignored a significant number of papers that were shown to them, leaving some completely uncited even after they had been displayed dozens of times. This was not a case of the machines making up fake references; every paper they cited was real, and every paper they ignored was real. The failure was not in hallucination, but in a shared, invisible filter that decided which real papers were worth reading.
Perhaps the most unsettling discovery was that this filter was the same for all the models, regardless of which company built them. Whether the AI was made by OpenAI, Google, or Anthropic, they all favored the exact same papers and ignored the exact same ones. The researchers found that the preferences of these eleven different models were so similar that they could be described as a single, shared map of what a "citable" paper looks like. This map was not based on the authors' names or the journal they appeared in, since those were hidden. Instead, it was driven entirely by the text of the abstracts. The models seemed to agree on a specific style or type of content that deserved attention, favoring papers that sounded like core theoretical work while systematically overlooking applied research, even when the applied papers were just as real and relevant.
To understand if this was a learned memory of past citations or a genuine judgment of quality, the researchers tested the models again after rewriting the titles and abstracts of the papers. They changed the words completely, ensuring the models could not recognize the papers by their surface text, but they kept the underlying meaning intact. The models' choices did not change. They still picked the same papers based on the new wording, proving that their preference was rooted in the meaning of the content, not in a memorized list of famous titles. This suggests that the models have developed a convergent, shared intuition about what constitutes a good paper, one that is distinct from how human experts actually evaluate the same work.
The study also explored what happens when these models are used repeatedly over time. As the AI models wrote new papers and those new papers were added back into the pool for future selection, the original real papers became rarer. The researchers found that as the real papers became scarcer, the models cited them even more intensely, almost as if they were competing for the few remaining "good" options. However, this did not mean the models were paying more attention to the real literature as a whole; in fact, the total share of citations going to the real papers dropped significantly. The models were simply concentrating their limited attention on a shrinking group of survivors, creating a feedback loop where a fixed preference meets a diluting catalog.
This research demonstrates that the danger of artificial intelligence in science is not just that it might invent fake facts, but that it might silently agree on a narrow set of "correct" answers while ignoring the rest of the field. Even when every option is real and every paper is equally visible, these systems impose a common filter on scientific attention. Mixing different models together does not solve the problem, because they all share the same underlying bias. The solution, the authors suggest, requires changing the shared preference map itself, rather than just trying to diversify the sources of information. The machines are not just reading the literature; they are beginning to decide what the literature is, and they are all deciding on the same thing.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.