← Latest papers
🧬 biology

TDEP: Predicting Target Perturbation Response from Drug-Induced Transcriptomes

TDEP is a deep neural network that leverages protein language model embeddings and graph attention networks to accurately predict transcriptome-wide target perturbation responses from drug-induced profiles, matching the accuracy of structure-based methods while providing detailed gene-level mechanistic insights for target prioritization and drug repurposing.

Original authors: Jieun Sung, Sookyung Kim, Wankyu Kim

Published 2026-09-04
📖 7 min read🧠 Deep dive

Original authors: Jieun Sung, Sookyung Kim, Wankyu Kim

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Finding the right key for a biological lock is the central challenge of modern drug discovery. Scientists have long known that medicines work by attaching to specific proteins inside the body, much like a key fitting into a lock. If the key fits, it can turn the lock to stop a disease or start a healing process. For decades, researchers have tried to predict these matches using two main strategies. One approach looks at the physical shape of the molecules, trying to see if they fit together like puzzle pieces. The other looks at the chemical aftermath, observing how a cell's internal instructions change when a drug is introduced. While the shape-based methods are good at guessing if a match will happen, they often fail to explain what happens next. The chemical methods can show the aftermath but struggle to pinpoint exactly which protein caused the change, often getting lost in the noise of the cell's complex reactions.

A team of researchers at Ewha Womans University in South Korea has developed a new way to bridge this gap. They created a system called TDEP, which stands for Target-perturbed Differential Expression Profile. Instead of just predicting whether a drug will stick to a protein, TDEP predicts the specific ripple effect that protein will cause throughout the cell's entire instruction manual, known as the transcriptome. Imagine a cell as a vast library of books, where each book is a gene. When a drug hits a target protein, it doesn't just silence one book; it changes the reading order of thousands of others. TDEP learns to read the library's reaction to a drug and then works backward to predict exactly which protein was the original cause of that disturbance. This allows scientists to not only find the match but also understand the biological story that follows.

The researchers built their system by teaching a computer to recognize patterns in massive amounts of biological data. They started with a huge collection of drug-induced changes in gene activity, gathered from a public database called the Connectivity Map. This database contains over a million snapshots of how different cells react to various chemicals. The team paired these chemical snapshots with the genetic sequences of the proteins the drugs are supposed to hit. They used a type of artificial intelligence that understands protein sequences, similar to how a language model understands sentences, to represent the target proteins. Then, they trained the system to connect a specific protein to the unique pattern of gene changes it produces. The system does not just say "yes" or "no" to an interaction; it generates a detailed map of how the cell's genes would turn up or down if that specific protein were blocked.

When the team tested their new system against existing methods, the results were striking. They compared TDEP to older models that rely on molecular shapes and others that rely on genetic data from experiments where genes are artificially silenced. In tests involving drugs the system had never seen before, TDEP achieved an AUC of 0.87 in correctly identifying interactions. This performance was comparable to the best shape-based models, which are currently considered the gold standard, and significantly better than the older genetic methods, which struggled to reach an AUC of 0.6. The researchers found that the system was particularly good at learning from the data it was given, avoiding the common pitfall of simply memorizing the answers. It learned to recognize the underlying biological logic that connects a protein to its cellular consequences.

One of the most important discoveries was how the system handles different types of unknowns. The researchers tested the model in two scenarios: one where the drugs were new but the proteins were familiar, and another where the proteins were entirely new. They found that the system uses two different tools to solve these problems. When the protein is known, the system relies on the detailed map of gene changes it has learned to predict the outcome. When the protein is new, the system leans on its deep understanding of the protein's genetic sequence to make an educated guess. This dual approach allows the system to be robust, working well whether it encounters a new drug or a new target. Furthermore, the system showed that it could transfer its knowledge across different types of cells. A model trained on one type of cancer cell could still make accurate predictions when applied to a different type of cancer cell, suggesting that the fundamental rules of how these proteins affect genes are shared across different biological contexts.

The team also looked closely at whether the system was actually learning the biology or just finding shortcuts. They examined the specific genes the system predicted would change when a target was blocked and compared them to real-world data from other experiments. They found a strong agreement. For example, when the system predicted the effect of blocking a specific heart-related protein, the genes it flagged as most likely to change matched the genes that actually changed in independent experiments where that protein was removed. This suggests that the system is capturing genuine biological signals rather than just statistical noise. The researchers noted that while the predictions are not perfect and the effect sizes are sometimes small, the ability to generate a gene-level explanation for a drug's action is a significant step forward. It provides a way to generate hypotheses about how a drug works, which can then be tested in the lab.

Despite these successes, the researchers are careful to point out the boundaries of their work. The system currently relies on data from a specific set of cancer cell lines and drugs that have already been profiled in the public database. It cannot yet predict the effects of completely new chemical structures that have never been tested in a lab. Additionally, the system predicts whether a drug will interact with a target, but it does not measure how strongly they bind. The researchers also acknowledge that the biological data they use comes from a specific technology, and the results might vary if applied to data from different measurement platforms. However, the core achievement remains: they have demonstrated that it is possible to predict the detailed transcriptional consequences of targeting a specific protein with high accuracy.

This work offers a new perspective on how we approach drug discovery. By moving beyond simple yes-or-no predictions, TDEP provides a window into the mechanism of action. It allows scientists to see the potential side effects or therapeutic benefits of a drug by looking at the specific pattern of gene changes it is predicted to cause. The researchers suggest that this approach could help prioritize which drug targets are worth pursuing, generate new ideas about how existing drugs might work, and even help find new uses for old medicines. As the databases of biological data continue to grow and the artificial intelligence models become more sophisticated, this method of linking protein targets to their full cellular consequences could become a standard tool in the quest to understand and treat disease. The study stands as a proof of concept that combining genetic sequences with cellular response data can yield insights that neither approach could achieve alone.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →