← Latest papers
🤖 machine learning

NRCD: An Open Database of Collegiate Running with Unified Performance Standardization

This paper introduces the National Running Club Database (NRCD), the first large-scale, open-access dataset of collegiate running performances in the United States, which includes over 128,000 results from 2004 to 2026 and provides a unified framework for standardizing race times to enable robust longitudinal, environmental, and gender-equity research.

Original authors: Jonathan A. Karr Jr., Ryan M. Fryer, Ben Darden, Nicholas Pell, Kayla Ambrose, Evan Hall, Ramzi K. Bualuan, Nitesh V. Chawla

Published 2026-08-18
📖 9 min read🧠 Deep dive

Original authors: Jonathan A. Karr Jr., Ryan M. Fryer, Ben Darden, Nicholas Pell, Kayla Ambrose, Evan Hall, Ramzi K. Bualuan, Nitesh V. Chawla

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of competitive running, performance is a conversation between the athlete and the environment. A runner's time is never just a measure of speed; it is a record of how that speed was achieved against a specific backdrop of distance, terrain, and weather. A race run on a flat, cool morning in a valley yields a different result than the same effort on a steep, hot hill in the mountains. For decades, scientists and coaches have understood that to truly compare runners, or to track how an individual improves over time, one must account for these external factors. Without this context, a fast time on a perfect day looks the same as a fast time on a grueling one, obscuring the true picture of athletic ability. Yet, despite the thousands of races held every year at American colleges, no single, organized collection of these results existed for researchers to study. The data was scattered across various websites, locked behind barriers that prevented large-scale analysis, leaving a vast gap in our understanding of how environmental conditions shape athletic performance.

A team of researchers has now filled that gap by building the National Running Club Database, a massive, open collection of race results that brings order to this chaotic landscape. The project aggregates nearly 129,000 approved race performances from over 28,000 athletes who competed in cross country, indoor track, outdoor track, and road races between 2004 and 2026. What makes this database unique is not just its size, but the way it treats the data. The researchers did not simply dump the raw numbers into a file; they applied a unified system to adjust every single time for the conditions under which it was run. This process, known as standardization, translates a time recorded on a long, hilly course in the heat into what that same runner would have achieved on a standard, flat course in ideal weather. By doing this, the team created a level playing field where a runner's performance can be compared fairly against their own past results or against the performances of others, regardless of where or when the race took place.

The database covers a wide spectrum of collegiate running, from elite Division I-caliber athletes to recreational club runners, ensuring that the findings reflect the full range of human endurance rather than just the top tier. Of the nearly 29,000 athletes included, about 36 percent are women, a significant step toward correcting a historical bias in sports research where male data has often dominated. The collection includes results from four distinct sports: cross country, which involves running over natural terrain; indoor and outdoor track, which takes place on controlled surfaces; and road races. For the most recent years, the database is incredibly rich in detail, capturing specific information about the course length, the amount of elevation gained and lost, and the weather conditions at the exact moment of the race. For older records, the data is still present but with fewer environmental details, reflecting the limitations of how the information was originally recorded.

To make sense of this vast amount of information, the researchers developed a step-by-step method to normalize the times. They started by adjusting for the distance, ensuring that a race run on a course that was slightly longer or shorter than the official distance was corrected to a standard length. Next, they accounted for the terrain, calculating how much slower a runner would naturally go on a hilly course compared to a flat one. They then factored in the altitude, recognizing that running at high elevations is harder due to thinner air, and adjusted the times to reflect sea-level conditions. Finally, they addressed the weather, specifically the heat. Using a mathematical model based on established coaching guidelines, they calculated how much slower a runner would be on a hot, humid day compared to a cool one. This heat adjustment is particularly important because temperature and humidity can drastically slow down endurance athletes, and without correcting for it, a runner might appear to have gotten slower over a season simply because the weather turned hot.

The results of this standardization process were striking. When the researchers looked at the raw, unadjusted times, the variation in performance for a single athlete from one race to the next was quite high. This noise made it difficult to tell if a runner was actually improving or just having a good or bad day due to external factors. However, once the times were adjusted for distance, elevation, and weather, the picture became much clearer. For female athletes, the variation in their performance from race to race dropped by more than half. For male athletes, the variation decreased by about one-third. This means that the standardized times are a much more reliable indicator of an athlete's true fitness level. The adjustments removed the "noise" of the environment, revealing the "signal" of the athlete's actual ability.

One of the most important findings from this work is that failing to account for these environmental factors can lead to misleading conclusions about how athletes improve. When the researchers looked at how much faster runners got over the course of a single season, they found that using only the raw times made it look like athletes were improving much more than they actually were. This is because the weather naturally cools down as the season progresses from late summer to autumn, making races faster regardless of the runner's fitness. By not adjusting for this cooling trend, previous analyses had inflated the perceived rate of improvement. The standardized data showed that while athletes do get faster, the magnitude of that improvement is smaller than raw numbers suggest. This distinction is crucial for coaches and scientists who want to understand the true effects of training without the distortion of changing weather.

The database also highlights the importance of looking at men and women separately. The researchers found that the adjustments affected the two groups differently, largely because the standard race distances for men and women in cross country are different. The process of converting times to a standard distance had a larger impact on the variation seen in women's results than in men's. This reinforces the idea that gender-specific analysis is necessary; combining the data into a single group would obscure these important differences and lead to inaccurate models. The team explicitly recommends that any future research using this data should treat men and women as distinct groups to ensure the findings are valid.

Beyond the numbers, the project represents a shift in how sports data is collected and shared. Instead of relying on small, hand-picked samples of race results, which often skewed toward male athletes, this database is built on a community-driven model. Coaches, athletes, and volunteers submit race results, which are then reviewed by experts to ensure accuracy before being added to the public record. This approach has allowed the database to grow steadily, adding hundreds of new meets each year. The team has also made the tools they used to standardize the data available to everyone through a free software package. This means that other researchers, coaches, and even athletes can apply the same rigorous adjustments to their own data, allowing for fair comparisons across different regions and conditions.

The implications of this work extend far beyond the track and field. By providing a clear, standardized view of performance, the database opens the door for new types of research into how environmental stressors affect the human body. Scientists can now study how heat, altitude, and terrain influence endurance in a way that was previously impossible due to data limitations. This could lead to better training strategies, improved safety guidelines for racing in extreme conditions, and a deeper understanding of the physiological limits of human performance. Furthermore, by ensuring that women's data is included and analyzed with the same rigor as men's, the project supports a more equitable approach to sports science, addressing long-standing gaps in our knowledge of female athletic performance.

The researchers are careful to note what their data cannot tell us. While they have adjusted for the environment, they cannot account for the internal state of the athlete on race day. Factors such as sleep quality, nutrition, stress from schoolwork, minor injuries, or even the psychological pressure of a big race are not captured in the results. Two runners with the same standardized time might have achieved it under very different personal circumstances. The database also does not include information about training loads, meaning it cannot explain why an athlete is fast or slow, only how fast they were under specific conditions. These limitations are acknowledged and serve as a guide for future studies, suggesting that the next step in this field will be to combine these standardized race results with data on training and health to get a complete picture of the athlete.

In the end, the National Running Club Database offers a new lens through which to view the sport of running. It transforms a collection of disparate race times into a coherent story about human endurance, stripped of the confounding variables of weather and terrain. It shows that when we remove the noise of the environment, the true patterns of athletic performance emerge with greater clarity. This work does not just provide a larger dataset; it provides a better way of thinking about the data itself, demonstrating that fairness in comparison requires more than just looking at the clock. It requires understanding the world in which the clock was running. As this resource continues to grow and as more researchers begin to use these tools, it promises to deepen our understanding of what it means to run fast, and to do so in a way that honors the experiences of all athletes, regardless of their gender or the conditions they face.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →