← Latest papers
💻 computer science

Semantic Alignment, Chronological Distance, and Citation-Network Prominence in Citation Formation: Evidence from Ten Citation Corpora

By applying a rigorously leakage-mitigated framework across ten bibliographic corpora, this study demonstrates that chronological distance is the dominant predictor of citation formation, consistently outweighing the predictive power of semantic alignment and citation-network prominence under temporally strict conditions.

Original authors: Moses Boudourides

Published 2026-07-17
📖 6 min read🧠 Deep dive

Original authors: Moses Boudourides

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Great Library of Tomorrow

Imagine the entire history of human knowledge as a massive, ever-growing library where every new book is a scientific paper. When a new book is written, the author must decide which older books to mention in their footnotes. These footnotes are called citations, and they are the lifeblood of science: they show who built whose ideas, how fast new discoveries are spreading, and which topics are the "hot" ones right now. For decades, scientists have tried to build a crystal ball—a computer program that can look at a new paper and guess exactly which older books it will cite. This is called citation prediction.

But here's the tricky part: time travel is impossible. In the real world, when an author writes a paper in 2020, they can only see books published before 2020. They cannot see the future. However, many computer programs trying to predict citations have been cheating. They peek at the "future" by looking at the entire library as it exists today, including books that weren't even written yet. It's like trying to guess who will win a soccer match by looking at the final scoreboard before the game starts. This "peeking" makes the computer seem super smart, but it doesn't tell us how real humans actually make decisions. This paper asks a simple but tough question: If we stop the computer from peeking into the future, what actually drives a scientist to cite one paper over another? Is it because the ideas match perfectly? Is it because the paper is famous? Or is it simply because the paper is new?

The Time-Travel-Free Experiment

The author of this paper, Moses Boudourides, decided to build a "time-travel-free" laboratory to find out. He gathered ten huge collections of scientific papers from five different worlds: Science (like protein folding and CRISPR), Engineering, Biomedicine, Social Science, and the Humanities (like Film Studies). He then built a strict set of rules for his computer program: it could only use information that was available before the citing paper was published. No future data, no peeking at how famous a paper became later, and no using metadata that was updated after the fact.

Once the computer was forced to play fair, the results were surprising. The computer wasn't perfect—it got the right answer somewhere between 53% and 73% of the time, depending on the field. But the real story wasn't about how well it guessed; it was about why it guessed what it did. The author broke the reasons for citing a paper down into three buckets: Semantic Alignment (do the ideas match?), Chronological Distance (how old is the paper?), and Network Prominence (how famous is the paper?).

The biggest shocker? Time wins, almost every single time.

In almost every field studied, the most powerful predictor of whether a paper would be cited was simply how recent it was. If you have two papers that are about the same topic, the computer was far more likely to pick the one published just last year than the one from ten years ago, even if the older one had a slightly better match in its text. In fact, once the computer knew how old the paper was, knowing the exact words in the title or abstract didn't help much at all. In four out of the ten fields, the papers that actually got cited were even less semantically similar to the new paper than the ones that didn't get cited. It's as if the computer was saying, "I don't care if the ideas are a perfect match; I just want the newest thing on the shelf."

However, there were two important exceptions where time did not rule the game. In the field of Memory Studies (Humanities), the ideas (semantic alignment) mattered more than the age. This makes sense because in fields like film or history, researchers often reach back to old, classic ideas that are still relevant, rather than just chasing the latest news. And in the Social Science field of Income Inequality, time didn't matter at all; the age of the paper provided no signal for prediction. Instead, in that specific field, a paper's "fame" (Network Prominence) was the primary driver.

The "Famous" Paper Myth

You might think that famous papers—those with thousands of citations—would be the ones everyone copies. The paper calls this "Network Prominence." But when the computer was forced to ignore the future and only look at what was known at the time, being famous didn't add much extra power in most fields. Why? Because in these tight-knit groups of researchers, the "famous" papers were usually just the older papers that had been around long enough to get noticed. Once the computer knew the paper was old (or new), knowing it was famous didn't tell it anything new. The "fame" was just a side effect of time.

What This Means for the Future

The paper concludes that for most scientific fields, when a paper is published matters more than what it says, at least when you are looking at a specific group of researchers working on the same topic. The "Matthew Effect"—the idea that the rich get richer and famous papers get more famous—is still true, but it's mostly because those papers have had more time to accumulate citations, not because they are inherently better.

The author is careful to say this isn't a magic formula for predicting the future of science. The computer still couldn't predict citations perfectly, and the results depend heavily on how you define the group of papers you are looking at. However, this study proves that if we want to understand how science really works, we have to stop letting our computers cheat by looking at the future. When we do, we see that science is driven less by perfect idea-matching and more by the relentless march of time. The newest ideas are the ones that get the most attention, and the "famous" papers are often just the ones that have been around long enough to be seen. But in some fields, like Memory Studies, the depth of the idea still trumps the date, and in others, like Income Inequality, the network of fame matters more than the clock.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →