Development of Longitudinal, Linked Maternal-Infant Cohorts using the Epic Cosmos Electronic Health Record Dataset
This study establishes a rigorous, reproducible framework for creating longitudinal, linked maternal-infant cohorts within the Epic Cosmos dataset, demonstrating that the resulting data is highly complete, representative of the US birth population, and suitable for epidemiologic research despite significant attrition when enforcing strict longitudinal follow-up criteria.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Hospitals today generate a vast, continuous stream of digital records for every patient they treat. These electronic health records capture everything from a routine blood pressure check to the specific medications a doctor prescribes, creating a detailed, real-time map of a person's health journey. For researchers studying pregnancy and newborns, this data holds immense promise. It offers a way to look at millions of cases at once, far beyond what is possible with small, local studies. However, a major hurdle has always been connecting the dots. In many systems, a mother's records and her baby's records live in separate silos, making it difficult to study how the health of one affects the other over time. Furthermore, researchers have often worried that data from electronic records might be incomplete or only represent a specific type of hospital, rather than the diverse population of the entire country.
A team of researchers set out to solve these problems using a massive, centralized database called Epic Cosmos. This system gathers structured health information from hundreds of hospitals and clinics across the United States and a few other locations, creating a single, unified view of patient care. The researchers wanted to prove that they could build a reliable, long-term study group by linking mothers and their babies within this digital network. They needed to show that the data was complete enough to trust, that the people in the database looked like the general population, and that they could follow these families from the first prenatal visit through the baby's early months without losing track of them.
To do this, the team started with nearly three million pregnancies that ended in a live birth between 2023 and 2024. They began by linking the record of every newborn to their mother's record, creating a paired dataset. This initial step required that the baby was born at a hospital using the system and that the mother's file contained the necessary family history link. This process successfully connected the vast majority of the data, resulting in a cohort of almost three million pregnancies. The researchers then tested how well the data held up when they added stricter rules to ensure they had a continuous story for each family. They required that the mother had at least one doctor's visit in the first three months of pregnancy, another in the second three months, a check-up after the baby was born, and a follow-up visit for the baby within three months of leaving the hospital.
Each time the researchers added a new requirement to ensure a complete medical story, the number of families in the study group shrank. By the time they required all four types of visits, the group had reduced to about 587,000 pregnancies. This drop was expected, as it excluded families who received care at different hospitals or who missed certain appointments. Crucially, the researchers found that the remaining group still looked very much like the original, larger group. The mix of ages, races, and geographic locations remained consistent with national birth records. The data was also remarkably complete; for most key details like the baby's birth weight, the mother's age, and the type of birth, the information was missing less than one percent of the time. Even for variables that are often harder to capture, such as body mass index at delivery and race, the data was available for more than three-quarters of the cases.
To prove that this linked data could answer real medical questions, the team conducted a specific test. They looked at women who already had high blood pressure before or early in their pregnancy to see if their blood pressure readings in the first trimester predicted a serious complication called preeclampsia later on. Preeclampsia is a condition marked by dangerously high blood pressure and organ damage that can occur during pregnancy. The analysis showed that among women with chronic high blood pressure, those whose systolic blood pressure—the top number in a reading—was 140 or higher in the first trimester were significantly more likely to develop preeclampsia later in the pregnancy. This finding confirmed that the digital records were accurate enough to detect subtle but important health patterns.
The study concluded that Epic Cosmos is a powerful tool for understanding pregnancy and newborn health on a national scale. The researchers demonstrated that they could create a long-term, linked view of mothers and babies that is both large and representative of the United States. While the process of requiring continuous care visits does reduce the total number of families available for study, it does not distort the characteristics of the group, meaning the results remain trustworthy. The work highlights that while electronic health records are a rich resource, researchers must be mindful of how they select their study groups to avoid missing the most vulnerable patients who may not have continuous care in the system. Ultimately, this approach opens the door for future studies to explore complex health questions using real-world data from millions of people.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.