AttribRank: An Open and Reproducible Framework for Diagnosing Attribution Sensitivity in Author-Level Research Assessment
This study introduces AttribRank, an open-source Python framework that quantifies how author-level research rankings shift based on evaluation choices like collaboration rules and indicator types, revealing that these variations are structured by disciplinary fields rather than random noise.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of modern science, how do we decide who is the most important researcher? For decades, the answer has often come down to numbers: how many papers a scientist has written, and how many times other scientists have cited those papers. These counts are used to make life-altering decisions about hiring, promotions, and funding. But these numbers are not simple facts like the weight of a rock; they are constructed summaries that depend heavily on the rules used to calculate them. If you change the rule for how credit is shared among a team of authors, or if you decide to ignore citations a scientist gives to their own work, the final score can shift. This means a researcher's standing is not a fixed point on a map, but a position that moves depending on the lens through which it is viewed. The question facing the scientific community is not just who is at the top, but how much that ranking changes when the rules of the game are tweaked.
To answer this, a team of researchers has built a new tool called AttribRank, designed to measure exactly how much a scientist's ranking wobbles when the evaluation rules change. Imagine a scientist standing on a ladder; the researchers wanted to know how many rungs they would slide up or down if the definition of "climbing" changed slightly. They did not invent new ways to measure success, but rather took four common ways of adjusting the score—how credit is split in a team, how self-citations are handled, how author order is weighted, and which type of score is used—and applied them to a massive group of highly cited scientists. By running these different rule sets through their software, they could see how much the relative standing of each scientist shifted. The result was a clear picture of how fragile or stable these rankings really are.
The researchers applied their framework to a database containing over 230,000 highly cited scientists from around the world, covering a vast array of fields from medicine to physics. They found that the rankings are indeed sensitive to the rules used. On average, a scientist's position shifted by about ten percentile points when the rules changed. To put this in perspective, if a researcher was in the top 10 percent of their field under one set of rules, they might drop to the top 20 percent under another, or rise even higher. This is not a tiny fluctuation; it represents a non-trivial change in standing that could influence real-world decisions. The study showed that these shifts are not random noise or errors in the data, but structured consequences of the choices made by the people doing the evaluating.
Not all the rules caused the same amount of movement. The researchers discovered that excluding self-citations—removing citations a scientist gives to their own previous work—caused the smallest amount of change in rankings. In contrast, the biggest shifts occurred when the researchers switched from a standard citation count to a more complex, composite score that combines several different metrics. Changing how credit is shared among co-authors or how the order of authors on a paper is interpreted also caused significant movement, though slightly less than switching the entire scoring system. This tells us that the choice of which overall score to use matters more than the decision to filter out self-citations when it comes to reshuffling the leaderboard.
The study also revealed that these shifts are not the same for everyone. The amount a ranking moves depends heavily on the specific scientific field a researcher works in. For example, in fields like physics and astronomy, where large teams are common, the way credit is shared among authors had a large impact on who ended up at the top. In other fields, the order of authors on a paper was the most influential factor. This means that a single rule cannot be applied fairly to every discipline without considering how scientists in that field actually work. The researchers found that the specific sub-field a scientist belongs to is often a better predictor of how sensitive their ranking will be than their country of origin.
When the researchers looked at national differences, they found that once they accounted for the types of fields scientists worked in and how long they had been publishing, the country itself explained very little of the variation in ranking sensitivity. While some countries showed slightly higher or lower average shifts, these differences were small compared to the differences seen between individual scientists and between different scientific disciplines. This suggests that national research systems are not the primary drivers of these ranking changes; rather, the structure of the science itself and the specific choices made in the evaluation process are far more important.
The researchers emphasize that their tool, AttribRank, is not a new way to rank scientists or a measure of who is truly the best. Instead, it is a diagnostic tool, like a stress test for an evaluation system. It does not tell you who should get a job or a grant; it tells you how much that decision would change if you used a slightly different set of rules. By making these shifts visible and measurable, the tool allows universities, funding agencies, and policymakers to see the assumptions hidden behind a single number. It encourages a more transparent approach where the rules of the game are acknowledged and discussed, rather than treating a ranking as an unchangeable fact. The study concludes that for research assessment to be responsible and fair, we must stop looking for a single, perfect number and start understanding how our choices shape the results we see.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.