Building a clinical and academic health research network using data science and artificial intelligence
This study demonstrates that using data science and artificial intelligence to systematically analyze publication metadata and construct graph-based models is a scalable, efficient, and unbiased alternative to manual methods for building clinical and academic health research networks focused on alcohol, addiction, and mental health in Northern England.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of medical science, solving complex health problems rarely happens in isolation. When researchers, doctors, and institutions join forces, they create a research network—a structured way to share knowledge, resources, and patients to tackle issues that are too big for any single person to handle alone. These alliances are vital for turning scientific discoveries into real-world treatments and policies. However, building these networks has traditionally been a slow, manual process. Scientists often rely on personal contacts, surveys, or institutional records to find who is working on what. This approach is labor-intensive, prone to human error, and quickly becomes outdated as new experts emerge and old teams dissolve. It is like trying to map a vast, shifting landscape by walking every inch of it on foot; it works, but it is exhausting and you might miss the most important paths.
A team of researchers from the University of Hull has proposed a different way to build these connections, one that uses the speed and pattern-recognition power of computers. They asked whether data science and artificial intelligence could automatically find the right people to join a research network, specifically for those studying alcohol, addiction, and mental health in Northern England. Instead of relying on phone calls or spreadsheets, they built a system that scans millions of scientific articles to find the authors, checks their identities, and maps out who is already working together. The goal was to see if a machine could do the heavy lifting of networking, creating a clear, up-to-date picture of the research landscape that is both scalable and systematic.
The researchers focused their attention on the northern part of England, an area facing significant public health challenges related to substance use and mental well-being. They wanted to identify the academics and clinicians in this region who were actively publishing on these topics between 2020 and 2025. To do this, they turned to PubMed, a massive, free database of medical literature. They selected twenty-six specific journals known for publishing work on alcohol, addiction, and mental health. Using a computer program, they searched these journals for every article published in the last five years. The software then pulled out the names of the authors and the universities or hospitals they worked for.
A major hurdle in this kind of work is that many researchers share the same names. To solve this, the team used a digital tool called ORCID, which acts like a unique ID card for scientists. The computer tried to match the names and affiliations found in the articles with these unique IDs to ensure that "John Smith" from one university was not confused with "John Smith" from another. Once the authors were identified and their identities confirmed, the team built a digital map of their relationships. In this map, every researcher is a point, and a line connects two points if they have written a paper together. This visual structure reveals who is collaborating with whom and how tightly knit the different groups are.
The results of this automated search were revealing. The system identified 420 unique researchers across twenty-one institutions in the North of England. When the team looked at where these researchers were based, they found that the Yorkshire and Humber region was the most active, producing 181 of the total publications. The North West followed with 162 publications, and the North East had 55. Among individual universities, the University of Sheffield stood out with 94 publications, followed by the University of York with 57 and the University of Liverpool with 52. The data showed a clear trend: while the total number of articles on these topics has seen a slight decline over the last few years, the concentration of expertise in the North remains significant.
When the researchers examined the connections between these scientists, they found a network that was moderately connected. On average, each researcher had collaborated with about seven other people. The map showed that while much of the collaboration happens within a single university, there is also a healthy amount of work happening between different institutions. The University of Sheffield emerged as a central hub in the Yorkshire and Humber region, acting as a focal point for many of the connections. However, the network was not a perfect web; the analysis suggested that only about four percent of all possible collaborations between these researchers actually exist, indicating that there is still plenty of room for new partnerships to form.
The study suggests that using artificial intelligence to build research networks is a feasible and powerful alternative to traditional methods. The authors argue that manual approaches are often too slow and inconsistent to be useful for large-scale projects, whereas their automated method can quickly generate a comprehensive list of potential collaborators. This approach offers a way to see the research landscape as it truly is, based on actual published work rather than who happens to know whom. However, the researchers are careful to note that their method is not perfect. Because the system relies on published articles, it might miss early-career researchers who have not yet published or clinicians who do not write for the selected journals. Additionally, the computer could not find a unique ID for every single author, so some names had to be used as they appeared, which carries a small risk of error.
Despite these limitations, the work demonstrates a clear path forward for how research communities can be formed and understood. By turning the vast, chaotic ocean of scientific literature into a structured map, this method provides funding agencies and policy makers with a tool to identify where expertise lies and where gaps might exist. The study concludes that this automated, data-driven approach is a substantial step up from the old ways of networking. It offers a way to systematically identify the people who are already doing the work, making it easier to bring them together to solve the complex health challenges facing Northern England and beyond. The future of this work will involve refining the tools to catch even more researchers and making the software available to others, so that building these vital connections becomes a standard part of the scientific process.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.