Scaling Electronic Health Record Foundation Models for Population Health Management
This paper introduces a scalable Electronic Health Record Foundation Model pre-trained on billions of cross-site medical events from over 5 million patients in the US and Taiwan, which demonstrates superior performance and generalization in predicting 11 chronic diseases compared to existing models by leveraging a unified code alignment framework and compute-optimal scaling.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Healthcare systems today often operate like a collection of isolated islands. A patient might see a specialist in one city, get a test done in another, and fill a prescription in a third, with each location keeping its own separate records. This fragmentation makes it difficult to see the full picture of a person's health over time. When doctors cannot easily connect these dots, they may miss the early warning signs of serious, long-term illnesses like heart disease or cancer. By the time these conditions are detected, they are often advanced, requiring expensive and intensive treatment. The goal of population health management is to solve this by looking at vast amounts of patient data to find people who are at risk before they ever feel sick, allowing for early, life-saving interventions.
To achieve this, researchers have turned to electronic health records, the digital logs of every visit, test, and medication a patient receives. However, these records are not written in a single, universal language. One hospital might use a specific code for a common procedure, while a hospital in a different country uses a completely different code for the exact same thing. This mismatch has made it nearly impossible to train computer systems to learn from data across different regions. A new study introduces a solution called EHR-FM, a massive computer model designed to read and understand these medical records regardless of where they come from. By learning from billions of medical events involving over five million patients from both the United States and Taiwan, the system has learned to spot the subtle patterns that precede chronic diseases, even when the data comes from two very different healthcare worlds.
The core challenge the researchers faced was not just the sheer volume of data, but the confusion caused by different coding systems. Imagine trying to learn a new language where the word for "heart attack" is spelled differently in every book you open. In medical records, a single drug or procedure might have hundreds of different names depending on the local billing system. To fix this, the team built a translator that converts all these different codes into a single, unified language before the computer ever sees them. They used a combination of official medical dictionaries and advanced artificial intelligence to map thousands of local codes to standard international ones. This allowed them to combine records from Taiwan and the United States into one massive, coherent dataset containing nearly ten billion medical events.
With this unified data, the researchers trained a large foundation model, a type of artificial intelligence that learns general rules from vast amounts of information. They tested how the model's performance changed as they increased its size and the amount of data it saw. They found that the model followed a predictable pattern: as they added more computing power and more data, the model became significantly better at its job. They built versions of the model ranging from very small to over two billion parameters, a measure of its complexity and capacity to learn. The largest version, trained on this cross-border data, proved to be exceptionally skilled at predicting the onset of major diseases.
When tested on real-world scenarios, the model demonstrated a remarkable ability to identify patients who would develop conditions like cancer, heart failure, or stroke within the next year. In the United States, the model achieved an AUC of 0.4 for future heart disease cases while keeping false alarms extremely low. In the Taiwanese data, it achieved an AUC of 0.7 for future cases with the same high level of accuracy. This high sensitivity is crucial for population health; it means the system can flag high-risk individuals for further screening without overwhelming doctors with false warnings. The model outperformed traditional computer programs that rely on simple rules and also surpassed other advanced AI models that had been trained only on data from a single location.
Perhaps the most striking finding was how well the model handled data from places it had never seen before. When tested on a completely different set of patient records from Stanford University, which used different coding systems and covered a different population, the model still performed better than previous state-of-the-art systems. This suggests that by learning from a diverse mix of data, the model developed a deeper understanding of human disease that transcends local differences. The researchers also discovered that simply copying the same data over and over to make a larger dataset did not help the model learn as well as adding new, diverse data from different regions. This indicates that the variety of the information is just as important as the quantity.
The study confirms that it is possible to build powerful, scalable tools for predicting chronic disease by unifying medical records from across the globe. The researchers showed that by aligning different coding systems, they could create a model that learns from the collective history of millions of patients, identifying risks that might otherwise go unnoticed. While the model is not yet a replacement for a doctor's judgment, it offers a powerful new way to prioritize care. By flagging high-risk individuals early, healthcare systems can shift from reacting to illness to preventing it, potentially saving lives and reducing the burden on medical resources. The researchers have made their code available to others, hoping that this approach will help build a future where proactive, data-driven care is accessible to everyone, regardless of where they live.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.