Towards understanding the disease landscape of clinical trials in Germany: Ontology and embedding-based pipelines versus Large Language Models for ICD-10 Harmonization
This study demonstrates that while a deterministic ontology and embedding-based pipeline can partially harmonize heterogeneous clinical trial condition data in Germany to ICD-10 standards, Large Language Models (specifically GPT-4o) significantly outperform it by achieving near-perfect agreement with expert human coding, thereby offering a superior solution for cross-registry disease landscape analysis.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine trying to organize a massive library where every book is written in a different language, uses a different alphabet, and sometimes just has a sticky note saying "something about a sick person" instead of a title. This is the current state of clinical trial data. Scientists run thousands of experiments to test new medicines, but they register these trials in different databases around the world. Some use strict medical codes, others use fancy medical dictionaries, and some just write it out in plain English or German. It's like trying to count how many people are studying "heart problems" when one database says "cardiac issues," another says "heart trouble," and a third just says "ouch, my chest hurts." Without a way to translate all these different descriptions into one common language, it's incredibly hard to see the big picture of what diseases scientists are actually studying and where the gaps might be.
To solve this, researchers need a translator. In the medical world, the "universal language" for diseases is called ICD-10, a giant list of codes that categorizes every known health condition. The challenge is taking the messy, mixed-up descriptions from trial registries and automatically turning them into the correct ICD-10 code. For a long time, scientists have tried to build computer programs that act like strict librarians, following a rigid rulebook to match words to codes. More recently, a new type of super-smart computer brain called a Large Language Model (LLM) has entered the chat. These models are trained on huge amounts of text and are great at understanding context and nuance, almost like a human who can read between the lines. The big question is: Can these AI brains do a better job of translating trial descriptions than the old, rule-based systems?
This paper dives right into that question by building a high-tech pipeline to translate clinical trial descriptions from four major registries in Germany into the standard ICD-10 codes. The researchers created a "deterministic pipeline," which is a fancy way of saying a multi-step robot process. First, the robot extracts the disease mentions. Then, it tries to match them using a strict medical dictionary (ontology) and a smart search engine that looks for similar meanings (embeddings). They even added a second version that re-ranks the top guesses to see if it can find the right one. But they didn't stop there. They also tested a powerful Large Language Model (GPT-4o), but with a twist: they forced the AI to follow the exact same step-by-step logic and rules that human experts use, just to see if the AI could be a better "librarian" than the robot.
The results were a bit of a shocker for the old-school approach. When they tested their rule-based robot pipeline against a "gold standard" of codes written by human medical experts, the robot got it right only about 49% of the time when looking for the specific three-letter code. It was decent at guessing the general category (like "heart disease" vs. "lung disease"), getting about 59% right, but it struggled with the details. The robot often got confused by similar-sounding diseases or couldn't find the right code in its dictionary.
However, the Large Language Model was in a completely different league. When the same AI was given the same rules and the same job, it achieved a staggering 96.7% accuracy. It agreed with the human experts almost perfectly. The paper suggests that the AI's superpower isn't just knowing the codes, but its ability to understand the context of a sentence. For example, if a trial mentions "adult-onset diabetes," the AI instantly knows that means Type 2 diabetes, whereas the robot might get stuck trying to match the exact words. The AI also managed to assign plausible codes to about 82% of the mentions that the robot gave up on and rejected.
The authors found that the robot's main failure wasn't in ranking the codes, but in even finding the right code to begin with. If the correct code wasn't in the robot's initial list of candidates, no amount of re-ranking could save it. The AI, on the other hand, was much better at figuring out what the right code should be, even when the text was messy or vague. The study concludes that while the old rule-based systems are okay for getting a rough idea of disease trends, Large Language Models are the clear winners for getting the details right. They suggest that future studies looking at the landscape of clinical trials should use these AI tools to get a much clearer, more accurate picture of what diseases are being studied, helping to ensure that research efforts are actually focused on the health needs that matter most.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.