Clinical trial success prediction from open registry text using classical machine learning
This study demonstrates that classical machine learning models, particularly random forests, can effectively predict clinical trial success using only open registry text from ClinicalTrials.gov, achieving moderate-to-strong performance across different trial phases and indications while mitigating data leakage through rigorous preprocessing.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Every year, pharmaceutical companies pour billions of dollars and decades of effort into developing new medicines. The path from a promising idea in a lab to a pill on a pharmacy shelf is long and fraught with failure. Most drug candidates never make it to approval, and when they fail after years of work, the financial loss is staggering. More importantly, the people who volunteer for these studies face risks without receiving the hoped-for benefits. Because resources are finite, every dollar and every hour spent on a drug that will eventually fail is a resource taken away from a drug that might succeed. The industry has long sought a way to spot these doomed candidates early, before too much time and money are invested, allowing researchers to stop the weak programs and focus on the strong ones.
To do this, scientists have tried to predict the outcome of clinical trials, which are the rigorous tests where new treatments are given to patients. These predictions are difficult because they rely on understanding complex human biology, the specific design of a study, and the behavior of a new drug. Traditionally, experts have used their experience to guess which trials will succeed, but this process is slow, subjective, and expensive. In recent years, researchers have turned to machine learning, using computers to find patterns in data that humans might miss. However, many of these computer models rely on private data that only big companies can access, or they require complex molecular information that isn't always available. This leaves smaller groups and academic researchers without a clear way to estimate the chances of a trial's success.
A new study by independent researcher Michael Doane explores whether the public records of these trials, which anyone can read online, contain enough information to make these predictions. The study focuses on a massive collection of trial records from a public website called ClinicalTrials.gov. This website acts as a global registry where researchers must register their studies before they begin, posting details about the disease being treated, the drug being tested, the number of patients involved, and the goals of the research. Doane's work asks a simple question: if you feed the text descriptions from these public records into a computer, can the computer learn to tell the difference between trials that will succeed and those that will fail?
To answer this, the researcher used a dataset called the Clinical Trial Outcome Dataset, which provides a label for about 125,000 past trials, marking them as either successful or unsuccessful based on their final results. The challenge was to build a model that could predict these outcomes using only the text available on the public registry at the time the trial was registered, without peeking at the final results or using any private company data. The researcher had to be extremely careful to avoid a common mistake called "data leakage," where a computer model accidentally learns the answer by reading a part of the text that was updated after the trial ended. For instance, a summary of a trial might be rewritten years later to include the final results, and if the computer reads that updated text, it isn't really predicting the future; it is just reading the past.
To prevent this, the researcher meticulously cleaned the data, removing any records where the text might have been altered after the outcome was known. They also excluded specific fields that directly described the results. Once the data was safe, they trained three different types of computer models to look for patterns in the remaining text. These models were tested on trials from different stages of development, known as phases. Phase 1 trials are the first time a drug is tested in humans to check for safety, Phase 2 trials look for signs that the drug actually works, and Phase 3 trials are large studies that confirm the drug's effectiveness and safety before it can be approved.
The results showed that the computer models could indeed find useful signals in the public text. The best-performing model, which used a method called random forest, was able to distinguish between successful and failed trials better than random chance. In the earliest stage of testing, Phase 1, the model performed quite well, correctly ranking the outcomes in about 74 percent of cases. This means that if you took a group of Phase 1 trials, the model could identify the ones that were more likely to succeed with a high degree of accuracy. The performance was lower for Phase 2 trials, which are notoriously difficult to predict because they depend heavily on whether the drug hits the right biological target, a detail often missing from public summaries. However, the model still performed better than chance in this stage as well.
One of the most interesting findings was that training the model on a wide variety of diseases actually helped it predict outcomes for a specific group of diseases known as neuroscience. Neuroscience trials, which deal with conditions like Alzheimer's disease and depression, have historically had very high failure rates. The researcher found that when the computer was trained on trials from all medical fields, not just neuroscience, it became better at predicting the success of neuroscience trials. This suggests that the general rules of how a clinical trial is designed and run are similar across different types of diseases, and learning from a broad range of examples helps the computer understand the specific challenges of the nervous system.
The study also tested the model against a set of 7,260 difficult cases that had been manually reviewed by human experts to ensure they were not easy to guess. Even against these tough examples, the model maintained its ability to distinguish between success and failure, proving that it was not just memorizing simple rules but was actually learning something about the nature of clinical trials. The researcher also checked to see if the model was relying on the names of the companies sponsoring the trials or the locations where they were held. When these details were removed, the model's performance dropped only slightly, indicating that the text describing the science and the study design was carrying most of the predictive weight.
This work demonstrates that public information, which has been available for years but largely unanalyzed in this way, holds significant value for the pharmaceutical industry. The models used in this study are relatively simple and do not require powerful supercomputers or secret data; they can run on a standard laptop. This makes the approach accessible to academic researchers, non-profit organizations, and smaller companies that do not have access to proprietary databases. By providing a way to estimate the likelihood of success using only open data, this method offers a new tool for screening drug candidates. It does not replace the need for expert judgment or the complex work of drug development, but it provides a reliable, objective starting point. It can help teams decide which programs deserve more attention and which might need to be stopped early, potentially saving time, money, and the well-being of the patients who volunteer for these critical studies.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.