The name-collision burden of name-based author attribution is unequal across name origins: a measurement in OpenAlex, and an identifier-based remedy (SigmaCV)
This study quantifies the significantly higher burden of name collisions faced by East-Asian researchers compared to their Anglophone and Other counterparts within the OpenAlex database, demonstrating that name-based attribution is inherently unequal and advocating for the adoption of identifier-based solutions like SigmaCV to ensure fair research assessment.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Technical Summary: The Name-Collision Burden in OpenAlex and the SigmaCV Remedy
Problem Statement
Responsible research assessment relies on the correct attribution of scholarly works to specific researchers. However, a significant portion of current practice still attributes works based on name strings rather than persistent identifiers. This approach fails systematically for common surnames and for names romanized from non-Latin scripts (particularly East-Asian names), leading to "name collisions" where multiple distinct author entities share the same normalized name. The paper posits that this ambiguity is not symmetric; the burden of name collision falls disproportionately on researchers with East-Asian names, creating a structural precondition for inequitable assessment. While prior work established the direction of this effect, this paper aims to quantify the magnitude of the burden across the entire population of ORCID-bearing authors in open scholarly data.
Methodology
The study utilizes a single, complete snapshot of the OpenAlex database from March 2026, comprising approximately 113.6 million author entities. The analysis focuses on the ~4.63 million authors within this snapshot who possess an ORCID iD.
Population Stratification: Authors are stratified by the country of their last-known institution into three groups:
- East-Asian: Japan, China, South Korea, Taiwan, Hong Kong.
- Anglophone: US, UK, Australia, Canada, New Zealand, Ireland.
- Other: Continental Europe and Brazil (serving as a coarse Latin-script baseline).
- Note: Country is used as a proxy for name origin and romanization exposure, not self-reported ethnicity.
Matching Operators: The authors define a transparent, self-defined exact-name operator (lowercasing, accent-folding, whitespace-collapsing) to compute collision counts for three distinct attribution strategies:
- Full-Name Entity Count (Upper Bound): The number of OpenAlex author entities sharing the exact normalized full name. This includes potential database over-splitting.
- ORCID-Restricted Count (Conservative Lower Bound): The subset of full-name matches that also possess an ORCID. Since distinct ORCIDs represent distinct real people, this serves as a conservative proxy for the number of distinct real individuals sharing a name.
- Initial + Surname Count (Legacy Worst Case): The number of entities sharing the same first-initial and surname, representing the worst-case scenario for name-string-only matching (e.g., legacy citation styles).
Statistical Analysis: The study computes exact population statistics (medians, interquartile ranges, 90th percentiles) rather than sampling estimates. It employs rank-biserial correlation () to measure effect sizes between strata and negative-binomial regression to adjust for publication volume, field, and career age.
Key Results
The study finds that the name-collision burden is large and sharply unequal across name origins, with the disparity widening as attribution strategies become less specific.
ORCID-Restricted Strategy (Real-Person Proxy):
- The median East-Asian researcher shares their name with 3 distinct real people (IQR 1–17).
- In contrast, the median Anglophone and "Other" researcher has a unique name (median 1).
- 10.5% of East-Asian researchers share their name with at least 100 distinct real people, compared to 0.9% for Anglophones and 0.1% for the "Other" group.
- The effect size is large ( vs. Anglophone; vs. Other).
Full-Name Entity Strategy:
- The median East-Asian researcher matches 9 entities (vs. 1 for others), reflecting both real collisions and database over-splitting.
Initial + Surname Strategy (Legacy Worst Case):
- The disparity is most extreme here. The median East-Asian researcher collides with 4,846 distinct identities (IQR 1,240–16,548).
- 92.3% of East-Asian researchers collide with at least 100 identities under this strategy.
- Anglophone and "Other" medians are 69 and 34, respectively.
- Effect sizes are very large ( and ).
Robustness Checks:
- Publication Volume: A negative-binomial regression adjusting for the number of publications shows that the East-Asian stratum still carries roughly 6.5 to 8.6 times the expected collision count of the Anglophone stratum. The gap is not an artifact of higher productivity among East-Asian authors.
- Sensitivity Analysis: Reclassifying authors by surname rather than institution country strengthens the observed gap, confirming that the country proxy is conservative.
Contributions
- Quantification of Name Ambiguity: The paper provides the first direct, full-population quantification of the name-collision burden in open scholarly data. It decomposes the burden into three layers (full-name entities, ORCID-restricted real people, and initial-based legacy matches), demonstrating that for romanized East-Asian names, initial-based attribution is not a minor nuisance but a categorical failure.
- SigmaCV: The authors present SigmaCV, an open-source (Apache-2.0), FAIR web application that assembles academic CVs by persistent identifier (ORCID) rather than by name. It serves as a deployed instance of the principle that attribution should be identifier-anchored to avoid name-based ambiguity.
Significance and Claims
The paper claims to measure name ambiguity, which is a necessary precondition for inequitable attribution, but explicitly distinguishes this from demonstrated assessment harm.
- Scope of Claim: The authors state they measure the exposure to risk (the collision burden), not the downstream consequences (e.g., specific citation distortions or evaluation outcomes). They argue that a large, unequally distributed burden makes inflated indicators and misattributed citations possible, thereby threatening the fairness of research assessment.
- Remedy: The paper advocates for a structural fix: shifting attribution from name strings to persistent identifiers. It emphasizes that this is a field consensus (implemented by OpenAlex, Scopus, Web of Science, and ORCID) rather than a novel invention, and that SigmaCV is one implementation of this principle.
- Limitations and Equity Trade-offs: The authors acknowledge that their population is restricted to ORCID holders, meaning the proposed remedy (identifier-based attribution) currently benefits those who already have identifiers. They note that ORCID adoption is uneven globally, so while identifier-anchoring solves the name-ambiguity problem for the identified, it may shift the equity gap to a "must hold an identifier" requirement. They do not claim this is a complete solution for all researchers but argue it is a necessary corrective for the identified population.
In summary, the paper provides empirical evidence that name-based attribution creates a systematically higher risk of conflation for East-Asian researchers compared to their Anglophone and European counterparts, motivating a shift toward identifier-based systems like SigmaCV to ensure fair research assessment.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.