Citation language changes little after findings fail to replicate: a longitudinal analysis
A longitudinal analysis of over 15,000 citing sentences reveals that subsequent papers rarely adjust their language to reflect failed replications or mention conflicting evidence, offering readers little indication of whether original findings have been successfully replicated.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Science is often imagined as a self-correcting machine, where new evidence automatically updates our understanding of the world. When a study is repeated and fails to produce the same result, the expectation is that the scientific community will notice, pause, and adjust how they talk about that original finding. This process relies on the language researchers use when they cite previous work. If a claim has been shown to be shaky, the sentences describing it in new papers should ideally become more cautious, or at least mention that the original result could not be confirmed. This is how knowledge is supposed to evolve: not by erasing the past, but by refining the description of what we know and what we do not.
A team of researchers set out to test whether this self-correction actually happens in the way we expect. They focused on a specific question: when a major, pre-planned attempt to repeat a famous study fails, do the authors of later papers change the words they use to describe that original finding? To find the answer, they looked at a large collection of scientific claims from psychology, economics, and social science. They identified 110 original studies that had been part of three major replication projects. For each of these studies, they gathered thousands of sentences from later papers that cited the original work. They then used advanced computer models to read these sentences and score them on a scale of certainty, measuring how definitively the authors stated the original claim. They compared the language used for findings that failed to replicate against the language used for findings that were successfully repeated.
The results revealed a striking lack of change. The researchers found that the wording used to describe failed findings remained almost identical to the wording used for successful ones. Even years after a replication attempt showed that an original result could not be reproduced, later papers continued to state those findings with the same level of confidence. The computer models detected no significant shift toward more cautious language. In fact, the difference in certainty between the two groups was so small it was statistically indistinguishable from zero. The study also looked for explicit mentions of the failure. They found that in the vast majority of cases, authors did not mention the failed replication at all. Only a tiny fraction of the sentences, roughly three to five out of every hundred, acknowledged that the original finding had been challenged or contradicted.
This pattern held true over time. The researchers tracked the language used in the eight to eleven years following the publication of the replication results. They found no gradual drift toward caution. The sentences describing the failed claims did not become softer or more qualified as time passed. The authors of the study suggest that this might be because researchers often reuse established phrases without re-reading the original evidence or checking for new, contradictory data. The original description of a finding can become a fixed part of the scientific record, passed along from paper to paper like a standard reference, regardless of whether the underlying evidence has been shaken.
The study does not claim that scientists are ignoring the evidence on purpose, but rather that the system of how findings are communicated has a kind of inertia. Even when a finding is shown to be unreliable, the way it is described in new research often remains unchanged. This has implications for how we read science and how we might build tools to help us understand it. If the language of science does not automatically reflect when a finding has failed, then readers, editors, and even artificial intelligence systems might be misled into thinking a claim is more solid than it actually is. The researchers conclude that for science to truly self-correct, we may need better ways to link original studies directly to their replication outcomes, ensuring that the story of a finding includes its failures, not just its initial success.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.