← Latest papers
💻 computer science

HonestAffinity: Leak-Aware Evaluation of Protein and Pocket Priors for Binding Affinity Prediction

The paper introduces HonestAffinity, a compact 1D predictor demonstrating that while protein embeddings and pocket markers improve performance on canonical splits, they degrade results on strict no-leak benchmarks, revealing a critical split-conditioned reversal that necessitates leak-aware evaluation protocols for reliable protein-ligand affinity prediction.

Original authors: Junhao Wei, Baili Lu, Zhenhong Peng, Wanyan Li, Zhirong Huang, Yanxiao Li, Yifu Zhao, Dexing Yao, Haochen Li, Xudong Ye, Sio-Kei Im, Yapeng Wang, Xu Yang

Published 2026-06-03
📖 5 min read🧠 Deep dive

Original authors: Junhao Wei, Baili Lu, Zhenhong Peng, Wanyan Li, Zhirong Huang, Yanxiao Li, Yifu Zhao, Dexing Yao, Haochen Li, Xudong Ye, Sio-Kei Im, Yapeng Wang, Xu Yang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to guess how strongly two puzzle pieces will snap together. One piece is a protein (a tiny machine in your body), and the other is a drug molecule. In the world of drug discovery, figuring out this "snap strength" (called binding affinity) is crucial. If you can predict it accurately, you can find new medicines faster.

For a long time, scientists have used two main ways to make these predictions:

  1. The 3D Approach: Looking at the actual 3D shape of the puzzle pieces. This is accurate but requires expensive, slow equipment to get the shapes.
  2. The 1D Approach: Just reading the "recipe" (the sequence of letters) for the protein and the drug. This is fast and cheap, but historically less accurate.

Recently, AI models using these "recipes" have gotten much better. But the authors of this paper, HonestAffinity, noticed a problem: The tests were rigged.

The "Leak" in the System

Imagine you are taking a math test. If the teacher gives you a practice test where the questions are almost identical to the real exam, you might get a perfect score. But that doesn't mean you actually understand math; it just means you memorized the answers.

In drug research, many AI models were being tested on datasets where the "practice" proteins looked very similar to the "exam" proteins. This is called a data leak. The models weren't learning to predict how new drugs work; they were just recognizing familiar patterns they had seen before.

The authors created a new, "leak-proof" test (called LP-PDBBind) where the exam proteins are completely different from the practice ones. They wanted to see if the fancy AI tricks still worked when the students couldn't cheat.

The Experiment: Three Different "Students"

The authors built a simple, fast AI model (HonestAffinity) and tested it in three different "outfits" to see which features actually helped:

  1. The "Super-Reader" (HonestAffinity-Pocket): This student reads the protein recipe using a massive, pre-trained encyclopedia (ESM-2) that knows a lot about biology, plus it gets a highlighter marking exactly where the drug might attach (the "pocket").
  2. The "Pocket-Only" Student (HonestAffinity-Pocket-NoESM): This student ignores the big encyclopedia and just learns the protein recipe from scratch, but still uses the highlighter for the pocket.
  3. The "No-Highlighter" Student (HonestAffinity-NoPocket): This student reads the recipe and uses the encyclopedia, but has no idea where the pocket is.

The Big Surprise: The "Reversal"

The results were counter-intuitive. It turned out that what helps you on a familiar test hurts you on a hard, new test.

  • On Familiar Tests (CASF-2016): When the exam proteins looked like the practice ones, the "Super-Reader" (with the big encyclopedia and the highlighter) was the best. It got the highest scores.
  • On the "Leak-Proof" Hard Tests: When the exam proteins were totally new and different, the "Super-Reader" actually got worse scores. The big encyclopedia and the highlighter confused the model.
    • Instead, the "Pocket-Only" Student (who ignored the big encyclopedia but kept the highlighter) became the champion on the hard tests.
    • Even better, if the student had no highlighter at all, they did surprisingly well on the hardest tests.

The Analogy:
Think of the "Big Encyclopedia" (ESM-2) like a student who memorized the history of one specific family.

  • If you ask them about that family, they are a genius.
  • If you ask them about a completely different family, they try to apply the old family's rules, get confused, and give the wrong answer.
  • The "Highlighter" (Pocket marker) is like a map. If the map is of the city you live in, it's great. If you are dropped in a new city and forced to use the old map, you will get lost.

The Conclusion: One Size Does Not Fit All

The paper argues that we shouldn't just pick one "best" AI model. The best tool depends entirely on the job:

  1. If you are studying a familiar target (one you've seen before) and you have a map of the pocket, use the Super-Reader. It uses all its knowledge to get the best score.
  2. If you are studying a brand-new, unfamiliar target (the "leak-proof" scenario), you should ditch the big encyclopedia. Use the model that learns from scratch, and if you have a map, keep it; if not, don't worry.

Why This Matters

The authors say that for a long time, the field has been reporting only the "familiar test" scores, making AI look better than it really is. They are calling for a new standard: Always test your model on both the "familiar" and the "strictly new" scenarios.

They also emphasize that their model is incredibly fast and cheap. It can train on a single computer in about 3 hours and make predictions in milliseconds. It's not the most powerful model in the world (3D models are still stronger if you have the data), but it is the most honest and practical tool for screening millions of new drug candidates when you don't have 3D shapes or when you are dealing with completely new biological targets.

In short: Don't trust a model just because it scores high on easy tests. Sometimes, the features that make it smart on familiar ground are the exact same features that make it fail when things get new and tricky.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →