Strategic Mapping of University Research Portfolios Using NLP and Machine Learning: An Integrated Analysis of Research Projects
This study presents an integrated machine learning framework that utilizes NLP, dimensionality reduction, and clustering algorithms to analyze a university's research portfolio, revealing that the Faculty of Engineering and Natural Sciences dominates the budget and identifying significant opportunities for strategic realignment alongside a 33.4% high-risk project profile.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Universities are engines of discovery, constantly generating thousands of research projects that aim to solve problems, invent new technologies, and expand human knowledge. To keep these engines running, institutions rely on research offices to manage the flow of ideas, money, and time. For decades, the way these offices evaluated their collections of projects—known as portfolios—relied heavily on counting finished products: how many patents were filed, how many papers were published, or how much money was brought in. This approach, however, looks backward. It tells administrators what has already happened but offers little insight into the current landscape of ongoing work, the hidden connections between different fields, or the potential risks lurking within a project before it fails. As the number of projects grows and research becomes more complex, the old method of manually reading hundreds of abstracts to find patterns has become too slow and prone to human bias. There is a growing need for a way to see the entire forest, not just the individual trees, using the tools of modern computing to understand the true shape of a university's research efforts.
In a recent study, researchers at Istanbul Sabahattin Zaim University tackled this challenge by applying advanced computer techniques to a massive collection of 631 scientific research projects conducted between 2015 and 2026. Instead of asking experts to read every single project description, the team used a form of artificial intelligence designed to understand language. They fed the titles and descriptions of these projects into a system that could read the text and convert the meaning of each sentence into a mathematical map. This process allowed the computer to "see" how similar or different two projects were based on their actual content, rather than just their assigned department or broad category. The researchers then used these maps to group the projects into natural clusters, revealing hidden themes that cut across traditional academic boundaries. They also built a system to spot projects that looked unusual or risky compared to the rest, creating a dynamic picture of the university's research health that goes far beyond simple spreadsheets.
The analysis began with a dataset of 631 projects, which was cleaned and refined to focus on 593 entries with sufficient detail to be meaningful. The researchers first stripped the text of unnecessary noise, such as common words that do not carry specific meaning, and then used a sophisticated language model to translate every project title into a dense set of numbers that captured its semantic essence. This step was crucial because it allowed the computer to understand that a project about "artificial intelligence in medicine" is conceptually closer to one about "machine learning for diagnosis" than to a project about "historical architecture," even if the words used were quite different. Once the text was converted into these numerical representations, the researchers applied dimensionality reduction techniques. Think of this as taking a complex, multi-layered object and flattening it onto a two-dimensional sheet of paper in a way that preserves the most important relationships between the parts. This made it possible to visualize the entire portfolio on a single screen, where the position of each dot represented a project's unique characteristics.
With the projects mapped out, the team used clustering algorithms to let the data speak for itself. These algorithms acted like a sorting machine, grouping projects that were close together in the mathematical space into distinct families. The researchers tested different ways of grouping the data and found that dividing the portfolio into ten distinct clusters provided the clearest and most stable picture. These ten groups were not arbitrary; they represented genuine thematic differences in the research being conducted. For example, one cluster was dominated by projects focused on materials science, nanotechnology, and advanced characterization, while another was filled with projects centered on artificial intelligence and data-driven engineering. A third group emerged around autonomous systems and defense technologies. By looking at the words that appeared most frequently in each group, the researchers could clearly label what each cluster was about, revealing that the university's research was not just a random mix of topics but a structured landscape of specialized domains.
The study also looked at the financial and operational side of these clusters to see if the themes matched the resources. They discovered that the Faculty of Engineering and Natural Sciences was the powerhouse of the portfolio, hosting nearly two-thirds of the total project budget. This concentration suggested a high level of productivity in technical fields but also pointed to a potential imbalance, as other faculties had significantly fewer resources. The researchers created a specific metric to measure how intensely money was being spent over time, dividing the total budget by the project duration. This revealed that some clusters were not only large in total size but also consumed resources at a much faster rate per month than others. The visual maps showed that the clusters with the highest budgets and the most intense spending were often the same ones that stood out as the most distinct in their thematic content, suggesting a strong link between the nature of the research and the scale of the investment required.
Perhaps the most critical finding of the study was the identification of risk. Using statistical methods designed to spot anomalies, the researchers analyzed the projects to see which ones deviated significantly from the norm. They found that roughly one-third of the entire portfolio fell into a high-risk category. These were not necessarily bad projects, but they were unusual in terms of their budget size, duration, or resource usage, making them more vulnerable to failure or management issues. The risk was not spread evenly across the university; instead, it was heavily concentrated in specific clusters. One particular group of projects had a 100 percent high-risk rating, meaning every single project in that cluster was flagged as unusual. Other clusters, including those focused on artificial intelligence and defense technologies, also showed significant concentrations of risk. This pattern suggested that certain types of research, particularly those involving large budgets and long timelines, carry inherent uncertainties that require closer monitoring. The study showed that these risks were not random; they were clustered in specific areas of the research landscape, allowing managers to target their attention where it was needed most.
The researchers validated their findings by testing the stability of their models. They ran the clustering process multiple times with different starting points to ensure that the groups they found were real and not just a fluke of the computer's random choices. The results were remarkably consistent, with the same projects ending up in the same groups every time. They also compared their language-based approach with older methods that simply counted how often words appeared. The newer method, which understood the meaning of the text, produced much clearer and more distinct groups, proving that understanding the context of the research was far more important than just counting keywords. This confirmed that the ten clusters they identified were a robust reflection of the university's actual research structure.
The implications of this work extend far beyond a single university. The study demonstrates that research portfolios can be managed with a level of precision and foresight that was previously impossible. By using these data-driven tools, technology transfer offices and research managers can move from a reactive stance, where they simply track what has already happened, to a proactive one. They can see emerging trends, identify interdisciplinary opportunities that might otherwise be missed, and spot potential problems before they escalate. The study suggests that the future of research management lies in integrating these analytical capabilities into daily operations, allowing institutions to make decisions based on a comprehensive view of their entire research ecosystem. The findings indicate that while the current portfolio has significant strengths, particularly in engineering and natural sciences, there is a clear need to support and balance research in other fields to ensure a more diverse and sustainable scientific future.
Ultimately, this research provides a new lens through which to view the complex world of academic inquiry. It shows that behind the thousands of individual project titles lies a coherent, structured, and sometimes risky landscape that can be mapped, understood, and managed. The study does not claim to have solved every problem in research management, but it offers a powerful, proven method for seeing the big picture. By turning text into data and data into insight, the researchers have provided a blueprint for how universities can better navigate the challenges of modern science, ensuring that their resources are directed toward the most promising and stable paths forward. The work stands as a testament to the power of combining human curiosity with computational power to reveal the hidden structures of knowledge.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.