← Latest papers
📄 medicine

LLM-assisted large-scale curation of peptide clinical trials maps trends in therapeutic development and trial outcomes

This study presents a validated, large language model-assisted workflow to curate and analyze 6,834 peptide clinical trials, revealing that while injectable delivery dominates and late-stage failures are less often due to toxicity compared to general drugs, overall phase-wise success rates remain modest.

Original authors: Emily Zhang, Uluc Birol, Maya Alev, Iris Caglayan, Emre Demirsoy, Mercan Deniz, Ali Salehi, Berke Ucar, Anat Yanai, Inanc Birol

Published 2026-08-28
📖 6 min read🧠 Deep dive

Original authors: Emily Zhang, Uluc Birol, Maya Alev, Iris Caglayan, Emre Demirsoy, Mercan Deniz, Ali Salehi, Berke Ucar, Anat Yanai, Inanc Birol

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

For decades, the world of medicine has relied heavily on two main types of drug molecules: small chemical compounds that fit into the body like keys in locks, and massive biological proteins that act like heavy machinery. Between these two extremes lies a middle ground of molecules called peptides. These are short chains of amino acids, the same building blocks that make up the proteins in our muscles and organs. Because they are built from nature's own materials, peptides can be incredibly precise, targeting specific diseases while often causing fewer unwanted reactions in the immune system. Since the first peptide drug, insulin, began treating diabetes in the 1920s, scientists have been eager to create more of them to fight cancer, metabolic disorders, and infections. However, turning these promising molecules into real medicines is notoriously difficult. Peptides often break down too quickly inside the body or struggle to pass through cell membranes to reach their targets. While researchers have developed many ways to fix these problems, such as changing the chemical structure or finding new ways to deliver the drug, it has been hard to see the big picture. The history of peptide drug development is scattered across thousands of separate clinical trials, buried in unstructured text, making it difficult to know which strategies actually work and which ones lead to dead ends.

A team of researchers set out to map this vast, unorganized landscape by using a new kind of digital tool. Instead of hiring armies of human experts to read through decades of trial records one by one, they trained an artificial intelligence system to do the heavy lifting. This system, a large language model, was taught to read the public records of clinical trials and extract specific, critical details: how the drug was delivered to the patient, the exact chemical sequence of the peptide, whether the trial succeeded or failed, and the specific reason for any failure. To ensure the machine was not just guessing, the researchers first tested it against a set of trials that had been carefully annotated by human volunteers. They found that the artificial intelligence agreed with the human experts almost as often as two different human experts agreed with each other. This validation gave them the confidence to let the machine read through a massive dataset of 6,834 peptide clinical trials spanning more than thirty years, a task that would have taken human teams years to complete.

The resulting map of peptide drug development reveals a field dominated by one specific delivery method. The vast majority of these trials, nearly 74 percent, relied on injections or infusions to get the drug into the body. This includes shots under the skin, into muscles, or directly into the bloodstream. While oral pills are the most common way people take medicine, they represented less than 15 percent of the peptide trials in this dataset. This heavy reliance on needles highlights a persistent challenge: peptides are fragile and often cannot survive the harsh environment of the digestive system. Despite the discomfort and inconvenience of injections, this method remains the most reliable way to ensure the drug reaches its target. The data also showed that while researchers are experimenting with different lengths and chemical charges for these peptide chains, there is no single "perfect" size or electrical charge that guarantees success. The trials that succeeded and those that failed showed a wide mix of these physical properties, suggesting that the answer lies in more complex biological features rather than simple measurements.

When the researchers looked at how often these trials succeeded, they found a pattern that mirrors the general drug development process but with some unique twists. In the earliest stage of testing, known as Phase 1, where the primary goal is to check if a drug is safe for humans, about 26 percent of peptide trials moved forward. This success rate climbed in the middle stage, or Phase 2, where efficacy is tested, reaching about 33 percent. The highest success rate appeared in Phase 3, the final large-scale test before a drug can be approved, where nearly 58 percent of trials reported positive outcomes. This is a notable finding because, in the broader world of drug development, the final stages are often where the most failures occur due to the high cost and complexity of testing on thousands of people. For peptides, however, the data suggests that if a candidate survives the early safety hurdles and reaches the final stage, it has a strong chance of proving effective.

Perhaps the most surprising insight from this analysis concerns why these trials fail. In the general world of pharmaceuticals, safety concerns and toxicity are major reasons for stopping a trial, especially in the later stages. However, for peptide drugs, safety was rarely the culprit. In the final Phase 3 trials, only about 4 percent of failures were due to the drug being toxic or unsafe. Instead, the overwhelming reason for stopping a peptide trial was that the drug simply did not work as intended. This lack of effectiveness accounted for more than a third of all failures. Another significant portion of trials ended not because of science, but because of business decisions, funding issues, or difficulties in finding enough volunteers to participate. The researchers also noted that trials with positive outcomes tended to involve much larger groups of people than those that failed, suggesting that having a robust study design and sufficient resources plays a crucial role in demonstrating a drug's true potential.

By organizing this chaotic history into a clear, structured dataset, the study offers a new way to understand the journey of peptide drugs. It confirms that while the path to approval is long and expensive, taking an average of several years for each phase, the final hurdles for peptides are less about safety and more about proving they work. The study also demonstrates that artificial intelligence can be a powerful partner in scientific research, capable of reading and understanding complex medical records with a level of accuracy that matches human experts. This approach allows scientists to look back at decades of data to identify patterns that were previously hidden, providing a clearer roadmap for the next generation of peptide medicines. As researchers continue to develop new ways to deliver these drugs without needles and to engineer more stable molecules, this large-scale view of the past offers a solid foundation for predicting which strategies will succeed in the future.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →