Bridging the Language Gap in Scholarly Data I: Enhancing Author Disambiguation Algorithms for Chinese Names
This paper proposes a rule-based, script-agnostic disambiguation framework that integrates co-authorship, citation, affiliation, and content data to effectively resolve author name ambiguities in both Romanized Pinyin and Chinese characters, achieving high F1-scores on a large-scale physics dataset from the China National Knowledge Infrastructure.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to organize a massive library where millions of books are written by people with the same name. In the West, if you see "John Smith," you might assume it's one person. But in China, names like "Wang Wei" are incredibly common—think of it like having thousands of people named "John Smith" in a single city.
The problem gets even trickier when these Chinese names are written for international databases. Instead of the original Chinese characters (like 王伟), they are often written in "Pinyin" (like Wang Wei). This is like taking a photo of a person and then describing them only by their shadow. Many different people can cast the exact same shadow, making it impossible to tell who is who just by looking at the name.
This paper is about building a smarter system to figure out which "Wang Wei" is actually the same person, even when the name looks identical to everyone else's.
The Problem: The "Shadow" Confusion
In the world of science, we need to know exactly who wrote which paper to give them credit. If we mix up two different "Wang Weis," we might think one person is a genius who wrote 1,000 papers, while the other is forgotten.
The authors found that existing computer programs are great at handling Western names but get confused by Chinese names, especially when they are just Romanized letters (Pinyin). They miss the subtle clues that Chinese characters provide.
The Solution: The "Detective's Toolkit"
The researchers built a new, rule-based "detective" algorithm. Instead of just looking at the name, their system asks four specific questions to decide if two papers were written by the same person:
- Do they have the same friends? (Co-authorship)
- Analogy: If two "Wang Weis" both hang out with the same group of scientists, they are likely the same person.
- Do they work in the same building? (Affiliation)
- Analogy: If both are listed at "Tsinghua University," that's a strong hint they are the same person.
- Do they talk about the same things? (Content Similarity)
- Analogy: This is the new superpower. Even if they don't know each other, if one "Wang Wei" writes about "quantum physics" and the other writes about "quantum physics," the system checks if the words in their papers are similar. It's like recognizing a person by their unique voice or handwriting style, even if you can't see their face.
- Do they quote each other? (Citations)
- Analogy: If one "Wang Wei" cites the work of another "Wang Wei," they are probably the same person giving themselves a shout-out.
If the answer to any of these questions is "Yes," the system merges the two names into one identity.
The Experiment: Testing the Detective
The team tested this system on a huge collection of 65,000 physics papers from China. They ran the test twice:
- Round 1: Using the original Chinese characters (the "face").
- Round 2: Using the Pinyin Romanization (the "shadow").
The Results:
- The new detective was incredibly accurate. It correctly identified the same person about 88-89% of the time.
- Crucially, it worked just as well on the "shadow" (Pinyin) as it did on the "face" (Chinese characters).
- It was much better than the old methods, especially at finding the people who should be grouped together (which the old methods often missed).
Why This Matters
Think of this like fixing a broken map. Before, if you tried to track a scientist's career across international databases, their path would look like a broken line with missing pieces because the computer couldn't tell the different "Wang Weis" apart.
This new method bridges the gap between the Chinese and English worlds. It ensures that a scientist gets credit for all their hard work, no matter how their name is written in a database. It stops the "ghost" scientists (who are actually many people merged into one) and the "split" scientists (who are one person split into many), making the global map of science fair and accurate for everyone.
In short: They built a smart filter that looks at who you know, where you work, and what you say, to solve the mystery of "Which Wang Wei is which?" even when the name alone is a dead end.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.