Mapping social determinants of health in NIH research funding with large language models
This study demonstrates that fine-tuned large language models, particularly ModernBERT-base, significantly outperform traditional keyword methods and zero-shot approaches in classifying social determinants of health within NIH grant narratives, while identifying threshold calibration as a critical strategy for addressing label imbalance.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Every day, a person's health is shaped by far more than just the medicine they receive or the doctors they visit. It is influenced by the air they breathe, the safety of their neighborhood, their ability to afford food, and whether they can access quality education. Scientists call these factors the social determinants of health. They are the invisible forces that drive why some communities thrive while others struggle, accounting for the vast majority of health outcomes. To fix these deep-rooted problems, governments need to know where their research money is going. In the United States, the National Institutes of Health acts as the world's largest public funder of health research, directing billions of dollars into specific projects. If this funding consistently ignores certain social problems, the scientific evidence needed to solve those problems never gets built.
For years, researchers trying to track how much money the government spends on these social issues have faced a difficult hurdle. They have had to read through thousands of grant applications, which are long, complex documents written by scientists. To find the relevant projects, they previously relied on simple computer searches that looked for specific keywords. If a grant proposal used the exact right words, it was counted. If it described the same problem using different language or wove the idea into a broader story, the computer missed it. This method was too rigid, like trying to find a needle in a haystack by only looking for needles that are exactly the same color. As a result, the true picture of where research money is flowing remained blurry, and gaps in funding for critical social issues went unnoticed.
A team of researchers recently decided to try a different approach, using a new generation of artificial intelligence known as large language models. These are computer systems trained on massive amounts of text that can understand context and nuance, much like a human reader. Instead of just hunting for keywords, these models can read a grant proposal and understand the underlying intent, even when the author does not use standard terminology. The researchers tested eleven different versions of these models to see which one could best identify mentions of social determinants in the grant applications. They wanted to know if these smart models could do a better job than the old keyword methods, and whether they could handle the messy reality of how scientists actually write about social problems.
The study focused on the "Public Health Relevance" sections across a dataset of over 2,000 grant applications from the 2023 fiscal year. The researchers first taught a model to recognize five specific categories of social determinants: economic stability, education access, health care access, neighborhood conditions, and social community context. They did this by having human experts carefully label a portion of the text, creating a set of correct answers for the model to learn from. They then tested the models to see how well they could find these categories in the remaining text. The challenge was significant because most grants do not mention these social factors at all, and when they do, the mentions are often rare and buried deep within the text.
The results revealed a surprising truth about how these artificial intelligence systems work. The researchers found that smaller models fine-tuned specifically on the downstream task outperformed the massive, general-purpose models that dominate headlines. In particular, ModernBERT-base achieved the highest performance when trained directly on the study's target classification task. It correctly identified the social factors in the text more often than the largest models, which were thousands of times bigger and relied on their general knowledge to guess the answers. The smaller models were also more stable and consistent, whereas the larger models sometimes struggled to learn the specific rules of the task.
However, the study also uncovered a critical lesson about how to use these tools. The researchers discovered that simply training the model was not enough; they had to adjust how the model made its final decisions. Because the social factors appeared so rarely in the text, the model tended to ignore them, assuming they were not there. By tweaking the model's sensitivity for each specific category, the researchers were able to recover a massive amount of missed information. This adjustment turned a system that was barely working into one that performed with high performance. Without this step, the most advanced models would have failed to capture the very data they were designed to find.
When the researchers looked closely at the different types of social factors, they found that the models behaved differently depending on the category. For clear, well-defined topics like economic stability or neighborhood conditions, the smaller, specialized models were excellent. They could spot these factors with high precision. But for more vague and complex topics, such as the social and community context, the larger, general models showed an advantage. These topics are often described in subtle ways that require a broad understanding of human relationships and society, a strength that the massive models possessed. This suggests that while smaller models are generally better and more efficient, the largest models still have a unique ability to understand the most abstract social concepts.
The implications of this work are significant for how science is funded and understood. The researchers demonstrated that it is possible to build a system that can scan decades of grant data to see exactly which social problems are being studied and which are being ignored. They showed that this can be done with relatively small, efficient computers that do not require expensive, massive hardware or secret, proprietary software. This makes it possible for public health officials and researchers to track funding trends in real time, ensuring that money is directed toward the areas where it is needed most. By using these tools, agencies can finally see the full landscape of their research investments, identifying blind spots where social determinants are underfunded and guiding future decisions to create a more equitable health system.
The study also highlighted a persistent gap in the current research landscape. Even with the best tools available, the analysis showed that certain areas, particularly education access and quality, remain significantly underrepresented in the grants that receive funding. This is not just a data error; it is a reflection of where the scientific community has chosen to focus its attention. When a social factor is consistently overlooked in research proposals, the evidence base needed to solve that problem remains thin. The ability to accurately map these trends means that policymakers can now see these gaps clearly and make informed choices to balance the portfolio of research, ensuring that the drivers of health disparities are addressed with the same rigor as the diseases they cause.
Ultimately, this research provides a new lens through which to view the machinery of scientific discovery. It moves beyond simple word counts to understand the true intent of the research being proposed. By combining the precision of specialized models with the broad understanding of large ones, and by carefully adjusting how these systems make decisions, scientists have created a powerful tool for transparency. This tool does not just count grants; it reveals the priorities of a nation's health research, offering a clear path toward a future where funding decisions are driven by a complete and accurate understanding of the social forces that shape our health.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.