← Latest papers
💻 bioinformatics

Biological rationales from language models enable leakage-resistant forecasts of target-indication success

The paper introduces PRIORITI, a leakage-resistant framework that leverages domain-instructed large language models to synthesize drug-agnostic biological evidence for generating calibrated, interpretable forecasts of target-indication success, outperforming existing baselines in both historical and prospective evaluations.

Original authors: Zhang, W., Xu, J., Zhang, Z., Liu, M., Sun, R., Al-lazikani, B., Shen, X., Kopetz, S., Wu, L., Zhao, B., Wu, C.

Published 2026-09-29
📖 5 min read🧠 Deep dive

Original authors: Zhang, W., Xu, J., Zhang, Z., Liu, M., Sun, R., Al-lazikani, B., Shen, X., Kopetz, S., Wu, L., Zhao, B., Wu, C.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

The journey from a biological idea to a medicine that saves lives is a long, expensive, and often frustrating road. Scientists can identify thousands of promising targets—specific genes or proteins that might fix a disease—and thousands of diseases that need fixing. But figuring out which combination of target and disease is strong enough to justify the billions of dollars and years of work required to build a drug is a massive guessing game. Historically, researchers have relied on human experts to sift through mountains of scattered data, looking for clues that a specific gene really causes a specific illness. This process is slow, uneven, and prone to missing the forest for the trees. The core challenge is not generating ideas, but filtering them: determining which biological hypotheses are solid enough to move forward before a single drug molecule is even designed.

A new study introduces a system called PRIORITI, designed to bring clarity to this chaotic early stage of drug discovery. The researchers built a framework that uses a large language model—a type of artificial intelligence trained on vast amounts of text—not to guess the outcome of a drug trial, but to act as a rigorous synthesizer of biological evidence. Instead of asking the AI to predict if a drug will work, the system asks it to read through all available human genetic data, animal studies, and biological pathways to construct a clear, drug-free argument for why a specific gene might treat a specific disease. The AI is strictly instructed to ignore any information about existing drugs, clinical trials, or regulatory approvals, ensuring it only looks at the raw biology. Once the AI writes this biological argument, a separate, traditional computer model analyzes the structure of that argument against a massive database of past drug successes and failures to calculate a probability of success.

The team tested this system on thousands of historical target-disease pairs, including a set of pairs where the outcomes were only known after the AI's training data cut-off date. This "out-of-time" test is crucial because it proves the system is predicting based on biological logic rather than simply memorizing past results. The results were striking. The PRIORITI system correctly ranked successful drug programs much higher than unsuccessful ones, achieving a level of accuracy that significantly outperformed both standard computer models that only look at gene names and disease definitions, and attempts to have the AI guess the outcome directly without the structured evidence step. The system was particularly good at identifying which hypotheses were likely to fail, correctly flagging the vast majority of unsuccessful pairs as low-probability candidates.

A key finding of the research is that the value comes from the structure of the biological argument, not just a simple score. The AI's ability to weave together different types of evidence—such as how well human genetics align with animal studies—created a rich description that the computer model could learn from. When the researchers tried to use the AI to guess the outcome directly, without this structured evidence step, the performance was poor and the confidence levels were unreliable. This suggests that the AI is best used as a translator that turns messy, scattered scientific facts into a coherent story, which a separate statistical model then uses to make a calibrated prediction. The system also provided a "rationale" for each prediction, explaining in plain language why a specific gene-disease pair was ranked high or low, making the process transparent and auditable for human scientists.

The study explicitly rules out the idea that the system's success was due to the AI simply remembering past drug trial results. The researchers verified that the AI rarely recognized the specific outcomes of the test cases from its training data, and even when it did, those cases were few and did not drive the overall results. They also demonstrated that the system works even when it encounters completely new genes or diseases it has never seen before, proving it learned to recognize patterns of biological plausibility rather than just memorizing specific names. However, the authors are careful to note that this system predicts the likelihood of a biological hypothesis being sound, not the success of a specific drug. A drug can fail for reasons unrelated to the biology, such as poor dosing or manufacturing issues, and a biologically sound idea can succeed through unexpected mechanisms.

Ultimately, this work offers a new way to prioritize the vast number of potential medicines before they enter the expensive clinical trial phase. By separating the task of gathering and summarizing evidence from the task of predicting the outcome, the researchers created a tool that is both powerful and transparent. It does not replace human judgment but provides a calibrated, biology-first probability that helps researchers decide where to focus their limited resources. The system suggests that the future of drug discovery lies not in asking artificial intelligence to be a fortune teller, but in using it to build a clear, evidence-based case for why a treatment might work, leaving the final decision to the rigorous testing of science.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →