← Latest papers
🧬 biology

EIHR-PROT: Explicitly Modelling of Homology Reliability for Protein Function Prediction

EIHR-PROT is a novel protein function prediction framework that outperforms existing methods by adaptively integrating ESM-2 sequence embeddings with homology-based priors through a learned gating mechanism that explicitly models the reliability of homology evidence.

Original authors: Oluranti Jonathan, Uwidia . E. Osalodion

Published 2026-08-21
📖 5 min read🧠 Deep dive

Original authors: Oluranti Jonathan, Uwidia . E. Osalodion

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Proteins are the molecular machines that keep life running, performing tasks that range from building cell structures to catalyzing the chemical reactions that power our bodies. To understand what a specific protein does, scientists often look at its amino acid sequence, the unique string of building blocks that makes it up. For decades, the most reliable way to guess a protein's function was to find another protein with a very similar sequence that had already been studied; if they looked alike, they likely did the same job. This method, known as homology, works beautifully when close relatives exist in the scientific record. However, as researchers explore the vast diversity of life, they increasingly encounter proteins that are so different from known ones that these traditional comparisons fail. In recent years, artificial intelligence models trained on millions of protein sequences have emerged as powerful tools, capable of spotting subtle patterns and evolutionary clues that simple sequence matching misses. Yet, a lingering question remains: when a new protein appears, should we trust the old method of finding a look-alike, or the new method of pattern recognition, or perhaps a combination of both?

A team of researchers at Covenant University in Nigeria has developed a new approach to answer this question, creating a system that does not simply choose one method over the other but learns when to trust each. They call their framework EIHR-PROT. Instead of forcing a fixed rule on how to combine the two types of evidence, their model acts like a smart filter that evaluates the quality of the "look-alike" matches for every single protein it examines. If the system finds a very strong, close match in its database, it leans heavily on that traditional evidence. If the matches are weak, distant, or non-existent, the system automatically shifts its confidence to the artificial intelligence model, which relies on the deep patterns it learned from studying millions of proteins. This dynamic adjustment allows the model to avoid the pitfalls of blindly trusting weak similarities while still capitalizing on them when they are reliable.

The researchers built their system using two main streams of information. The first stream comes from a sophisticated language model called ESM-2, which has been trained on a massive collection of protein sequences to understand their biological grammar. This model converts the sequence of a protein into a mathematical representation that captures its hidden structural and functional traits. The second stream comes from a standard search tool that looks for similar sequences in a database of known proteins. The innovation lies in how these two streams are joined. The researchers added a "gating" mechanism that examines specific clues about the quality of the database matches, such as how much of the protein sequence aligns and how strong the statistical score of the match is. Based on these clues, the gate decides exactly how much weight to give to the database match versus the AI prediction for that specific protein.

When the team tested this system on a large set of proteins with known functions, the results showed that this adaptive approach outperformed existing methods that use fixed rules for combining evidence. The model achieved high accuracy in predicting three major categories of protein function: what biological processes the protein is involved in, what molecular tasks it performs, and where it is located within a cell. In these tests, the system correctly identified the function of proteins with a high degree of precision, beating out other leading methods that mix sequence and homology data. The researchers found that the improvement was most noticeable when predicting biological processes, a complex area where the reliability of traditional matches can vary wildly. By explicitly modeling the reliability of the homology evidence, the system learned to ignore noisy or misleading matches that might have confused a simpler model.

To understand why this worked so well, the researchers ran experiments where they removed specific parts of their system. When they took away the ability to judge the reliability of the matches, the model's performance dropped, confirming that the smart gating mechanism was the key to its success. They also tested the system on a set of proteins that were very different from anything in the training database, simulating the discovery of entirely new biological entities. Even in these difficult cases, where traditional methods often struggle, the system maintained strong performance. This suggests that the model successfully learned to rely on the AI's deep understanding of protein patterns when the traditional "look-alike" evidence was too weak to be trusted. The study demonstrates that the future of protein function prediction does not necessarily require choosing between old and new methods, but rather building systems that can intelligently weigh the evidence available for each unique case.

The implications of this work extend beyond just getting better scores on a test. By making the decision-making process transparent, the system allows scientists to see exactly how much trust was placed in the database matches versus the AI predictions for any given protein. This interpretability is crucial for researchers who need to understand the basis of a prediction before acting on it. The authors note that while their current system handles protein sequences effectively, future versions could incorporate other types of biological data, such as protein structures or interaction networks, using the same flexible logic. For now, this research offers a clear path forward: a method that respects the value of traditional biological knowledge while fully embracing the power of modern artificial intelligence, adapting its strategy to the specific needs of every protein it encounters.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →