Real-world drug use in ATC and ICD-10: an expert-curated drug-diagnosis resource based on UK primary care and Danish hospitalization electronic health records
This paper presents a curated resource of 7,763 significant real-world drug-diagnosis associations derived from UK and Danish electronic health records, which was manually annotated by medical experts to reveal that most observed co-occurrences do not represent direct treatment and to highlight the necessity of accounting for inter-annotator variability.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Hospitals and doctors' offices generate a vast, continuous stream of records. Every time a patient receives a diagnosis or a prescription, a digital entry is made, creating a massive library of how medicine is practiced in the real world. To make sense of this information, researchers rely on standardized codes: one set of codes for diseases and another for medications. These systems act like a universal language, allowing computers to read and compare medical histories across different countries and decades. However, a simple match between a disease code and a drug code does not always tell the whole story. Just because a patient is taking a specific medication and has a specific diagnosis recorded at the same time does not mean the drug was prescribed to treat that specific illness. The medication might be managing a symptom, treating a side effect from another treatment, or addressing a completely different health issue that happens to exist alongside the main diagnosis. Understanding the true reason behind a prescription is crucial for researchers who want to use these records to discover new uses for old drugs or to study how diseases progress.
A team of researchers set out to build a clearer map of these relationships by looking at the actual records of approximately 1.5 million people in the United Kingdom and Denmark. They focused on two distinct types of healthcare data: records from general practitioners in the UK, which capture routine, everyday care, and records from hospitals in Denmark, which document more severe conditions and specialized treatments. By combining these sources, the team created a dataset containing nearly 736,000 pairs of drug and diagnosis codes that appeared together frequently. They then used statistical methods to filter this massive list down to about 7,763 pairs that showed a strong, significant link, meaning the drug and the diagnosis appeared together far more often than would happen by random chance.
The core of their work involved asking human experts to interpret these links. Six clinicians—three from the UK and three from Denmark—reviewed each of the 7,763 pairs independently. Their task was to decide if the drug was being used to directly treat the diagnosed condition, or if the connection was indirect. The experts labeled each pair as a direct treatment, an indirect relationship, or unsure. The results revealed a surprising reality about how medicine is used in practice. The vast majority of the strong statistical links the team found were not cases of a drug treating the specific disease it was paired with. Instead, most of these associations were indirect. For instance, a patient with cancer might be taking a drug to manage nausea caused by chemotherapy, or a patient with a chronic heart condition might be taking medication for a side effect of their primary treatment. In many cases, the drug was addressing a symptom or a secondary health problem rather than the main diagnosis recorded in the file.
The study also highlighted how different healthcare settings shape these records. The data from Danish hospitals showed a high frequency of drugs used for pain and anesthesia, as well as treatments for infections, reflecting the nature of hospital care where surgeries and acute illnesses are common. In contrast, the data from UK general practices was dominated by medications for respiratory issues, mental health, and cardiovascular conditions, which are typical of long-term management in the community. When the researchers looked at the overlap between the two countries, they found that the pairs appearing in both datasets were almost always clinically logical, such as drugs for heart disease paired with heart diagnoses. However, even among these clear cases, the human experts did not always agree. The level of agreement among the doctors varied, with some pairs causing significant debate. This disagreement was not seen as a failure but as a vital clue: it showed that real-world drug use is often complex and context-dependent, and that a single label cannot always capture the nuance of a clinical decision.
To make sense of these varying opinions, the researchers developed a scoring system that combined the doctors' judgments into a single confidence score for each drug-disease pair. This score did not just count votes; it accounted for the fact that some doctors were more cautious than others and that certain types of medical relationships are inherently harder to classify. The final resource is a curated list of thousands of drug-disease relationships, each tagged with a score that indicates how likely it is to be a direct treatment versus an indirect association. This tool allows other scientists to distinguish between a drug being used to cure a disease and a drug being used to manage the side effects of that disease or its treatment. By providing this level of detail, the study offers a more accurate foundation for using electronic health records in research, ensuring that future discoveries are built on a clear understanding of how medications are actually used in the real world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.