A human-validated procedure for linking PhD dissertation metadata to OpenAlex author profiles
This paper presents and validates a generalizable, human-verified method that combines character-level, topical, and transformer-based similarity measures to reliably link PhD dissertation metadata to OpenAlex author profiles, achieving high accuracy in tracking academic career trajectories and revealing that approximately 17.3% of North American PhD graduates cannot be linked to subsequent publications.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Every year, hundreds of thousands of students around the world complete their doctoral degrees, a milestone that marks their formal entry into the world of research. For decades, scientists who study how academic careers unfold have relied on a specific group of people to tell them what happens next: those who continue to publish papers in journals. By tracking these published works, researchers have built a picture of how long scientists stay active, how productive they are, and how their careers evolve. However, this picture has a significant blind spot. It only includes the people who kept publishing. It misses the vast number of graduates who leave the academic publishing world entirely, perhaps to work in industry, government, or teaching, or simply to stop writing for journals. Because these individuals do not appear in the standard databases used to track scientific output, they are effectively invisible to the studies that try to understand the full scope of a PhD's impact. Without knowing who is missing, any estimate of career success or attrition is based on a sample that has already been filtered by the requirement to keep publishing.
To solve this problem, a team of researchers has developed a new way to connect the dots between a doctoral dissertation and the subsequent career of its author, even if that author never published another paper. They focused on the massive gap between the record of a degree and the record of a career. The challenge is immense: a dissertation is often the only record of a person's work in a specific field, and the databases that hold millions of scientific articles are filled with people who share common names. A researcher named "John Smith" in a 2009 thesis could be the same person as a "J. Smith" who published a paper in 2015, or they could be two completely different people. The difficulty is compounded by the fact that many graduates change their names, use different initials, or publish under slightly different variations of their names. Previous methods often relied on simple name matching or on volunteers who had already linked their own work, both of which fail to capture the silent majority of graduates who do not continue in the publishing world.
The researchers tackled this by creating a computerized system that acts like a highly skilled librarian, comparing the details of a dissertation against millions of potential matches in a global database of scientific publications called OpenAlex. Instead of just looking at names, the system reads the titles of the dissertations and the titles of the published papers, breaking them down into small pieces of text to see how much they overlap. It also uses advanced technology to understand the underlying topics of the work, recognizing that a paper about "machine learning" might be related to a thesis titled "artificial intelligence" even if the words are different. The system generates a score that represents how likely it is that the person who wrote the thesis is the same person who wrote the articles. To make sure this system works, the researchers tested it by having human experts review thousands of potential matches, checking whether the computer's guesses were correct. They found that by combining the text matching with the topic analysis, the system could reliably identify the same person across different records with a high degree of accuracy.
The results of this new method reveal a startling reality about the academic workforce. When the researchers applied their system to a large group of PhD graduates from North America, they found that a significant portion of these individuals could not be linked to any subsequent publication in the database. Specifically, they estimate that between 10.2% and 20.7% of a typical cohort of PhD graduates never appear in the major bibliographic databases after receiving their degree, with a best guess of 17.3%. This means that nearly one in five doctoral graduates is missing from the standard maps of scientific careers. This group is not necessarily failing; they may be successful in other sectors, but their absence from the data skews our understanding of what happens after a PhD. The study also showed that the rate of "disappearing" from the publication record increases over time. While about 73% of graduates are still observed publishing in the year immediately following their degree, that number drops to roughly 50% by fifteen years later. This decline is even steeper when you include the graduates who never published at all, suggesting that the majority of researchers contribute to the scientific literature for only a limited window of time.
The researchers also tested whether modern artificial intelligence could replace their structured computer system. They found that while large language models could make very accurate judgments when they were highly confident, they were slower and slightly less effective at finding matches across the entire group than the specialized scoring system. The most efficient approach, they suggest, is to use the fast, specialized system to scan all the data and then use the artificial intelligence only to double-check the cases where the first system is unsure. This two-step process allows for a comprehensive view of the data without getting bogged down by the time and cost of checking every single record with a complex model.
Ultimately, this work does more than just fix a technical problem in data matching; it changes the story of the academic career. By accounting for the graduates who do not publish, the study provides a more honest and complete picture of the outcomes of doctoral training. It shows that the path after a PhD is far more diverse than the data has previously allowed us to see, with a substantial number of people moving into roles that do not involve publishing in academic journals. The method developed here is open and can be applied to other fields and countries, offering a new tool for policymakers and researchers to understand the true fate of the millions of people who earn a doctorate every year. It reminds us that the end of a dissertation is not always the beginning of a scientific career, and that the value of a PhD extends far beyond the pages of a journal.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.