← Latest papers
📄 medicine

A curated benchmark database for high-throughput mechanistic pharmacokinetic prediction

This paper introduces a curated, machine-readable database of 180 plasma concentration–time profiles across multiple species and administration routes to address the lack of public benchmarks, thereby enabling transparent, reproducible evaluation and improvement of high-throughput mechanistic pharmacokinetic prediction methods.

Original authors: Jeremy O. Jones, Rafał A. Bachorz, Michael S. Lawless, David W. Miller, Robert Fraczkiewicz, Viera Lukacova

Published 2026-08-19
📖 5 min read🧠 Deep dive

Original authors: Jeremy O. Jones, Rafał A. Bachorz, Michael S. Lawless, David W. Miller, Robert Fraczkiewicz, Viera Lukacova

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the early stages of creating a new medicine, scientists face a difficult balancing act. They must design molecules that can effectively fight a disease while also ensuring those molecules can travel through the human body to reach their target. This journey is known as pharmacokinetics. It involves how a drug is absorbed, how it moves through tissues, how the body breaks it down, and how it is eventually removed. For decades, this part of the process was often left until late in development because measuring it required expensive lab tests and complex animal studies. Consequently, researchers would design a drug based on how well it might work against a disease, only to discover later that the body could not absorb it or would eliminate it too quickly. To speed up discovery, scientists have turned to computer models that can predict these journeys before a single drop of liquid is tested. However, for these digital predictions to be trusted, they need to be checked against real-world data. Until now, there has been a lack of high-quality, organized data sets that allow different computer programs to be compared fairly, making it hard to know which tools are truly reliable.

A team of researchers at Simulations Plus, Inc., has addressed this gap by assembling a new, carefully curated database designed to test these prediction tools. They gathered information on 180 specific scenarios involving 76 different drug compounds tested in five different species: humans, monkeys, dogs, rats, and mice. For each scenario, they collected the full story of how the drug behaved over time, recording the concentration of the drug in the blood at every measured moment rather than just a single summary number. This approach captures the entire shape of the drug's journey, revealing details about how fast it is absorbed or how long it stays in the body. The team standardized the chemical structures of the compounds and the units of measurement, ensuring that every data point was consistent and ready for computer analysis. They also included information on how the drug was given, such as through an injection into a vein, a swallowed tablet, or a liquid solution, and noted the specific dose used in each case.

Using this new database, the researchers tested a high-speed computer system designed to predict these drug journeys. They ran simulations for every one of the 180 cases, feeding the system the chemical structure of the drug and the details of the study, such as the species and the dose. The system then generated a predicted curve showing what the drug concentration should look like over time. The researchers compared these predicted curves directly against the actual curves recorded in the real-world studies. They found that the computer model performed reasonably well, with about one-third of the predictions falling within a factor of two of the actual measurements. In the world of drug discovery, being within a factor of two is often considered a strong result, but the model performed even better when looking at a wider margin; nearly all of the predictions were within a factor of ten of the observed values. This level of accuracy suggests that the computer model can capture the general behavior of drugs across different animals and administration methods.

The study also revealed where the computer model struggles. The predictions were most accurate for drugs that are eliminated by the liver or kidneys, which are the primary pathways the model accounts for. However, the model was less accurate for certain types of drugs where the body absorbs them into the liver in a specific way that the current software does not fully simulate. This is a crucial finding because it tells developers exactly where to focus their improvements. Furthermore, the researchers tested whether the model could work using data from other sources, not just their own internal predictions. They found that the model could successfully run using input data from different computer programs, showing that the database is flexible enough to test various tools. This flexibility is vital because it allows different research groups to use the same benchmark to see how their specific methods compare, fostering a shared standard for improvement.

The database also includes a variety of complex molecules that break traditional rules of drug design, such as large cyclic peptides and macrocycles. These are newer types of medicines that are harder to predict, and the researchers found that their updated models handled these complex structures better than older versions did. By including these challenging cases, the database ensures that the tools being tested are robust enough for the future of medicine, not just the drugs of the past. The researchers acknowledge that the database is not perfect; it is limited to single doses and does not yet cover every possible way a drug might be administered, such as through an infusion or under the skin. They also noted that they had to combine data from different studies, which introduces some natural variation. However, they have made the database openly available and plan to keep it updated, inviting other scientists to report errors or add new data.

This work represents a shift in how drug discovery is evaluated. Instead of simply ranking computer programs on a leaderboard based on a single number, this approach encourages a deeper look at why a model succeeds or fails. By comparing the entire curve of a drug's journey, scientists can see if a model is failing during absorption, distribution, or elimination, providing a clear roadmap for fixing the software. The database serves as a shared resource that allows the scientific community to move beyond isolated experiments and toward a more transparent, reproducible standard. As the field moves toward using artificial intelligence to design medicines, having a reliable way to test these predictions is essential. This curated collection of real-world data provides that foundation, helping to ensure that the digital tools guiding the next generation of medicines are as accurate and trustworthy as possible.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →