A novel benchmark dataset for enzyme function prediction reveals the limitations of state-of-the-art models
The paper introduces EnzymARC, a novel benchmark dataset of structurally disrupted enzyme decoys, to demonstrate that current state-of-the-art enzyme function predictors rely heavily on phylogenetic shortcuts and fail to distinguish catalytically incompetent variants from functional enzymes, thereby highlighting the critical need for structure-aware negative examples in model training and evaluation.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
In the vast library of life, every living thing carries a set of instructions written in a code called DNA. Within these instructions are blueprints for tiny biological machines known as enzymes. These enzymes act as the workers of the cell, speeding up the chemical reactions that keep an organism alive, from digesting food to repairing damaged tissue. To keep track of the millions of different enzymes scientists have discovered, they use a standardized labeling system called Enzyme Commission numbers. Think of these numbers like a universal catalog system that tells a researcher exactly what job a specific enzyme performs. As scientists sequence the genomes of more and more organisms, they are finding new enzymes at a staggering rate. The challenge now is to figure out what these newly discovered enzymes do without having to test each one in a lab, a task that would take centuries. To solve this, researchers have turned to computers, training artificial intelligence to look at the sequence of an enzyme and predict its function based on patterns it has learned from known examples.
A team of researchers recently decided to test whether these computer models are truly understanding how enzymes work or if they are simply taking shortcuts. They created a new set of test data called EnzymARC to see if the best current programs could tell the difference between a working enzyme and a broken one. To build this test, they started with enzymes that were already known to work and then systematically damaged them. Using a computer model of the enzyme's 3D shape, they targeted the specific spot where the chemical reaction happens, known as the active site. They then altered the structure in a way that would destroy the enzyme's ability to function, creating what they call "decoy" sequences. These decoys were designed to look almost exactly like the original, working enzymes, except for the critical damage to the machinery that makes them work. They created these broken versions by changing the structure within specific distances from the active site, ranging from a very tight 5 Angstroms out to a wider 15 Angstroms radius, ensuring the rest of the enzyme remained largely unchanged.
The researchers then fed these broken decoys into three different types of computer prediction tools to see what the machines would say. One tool relied on finding similar sequences in a database, another used a method that learns by comparing different protein shapes, and the third was a deep learning model specifically trained to recognize non-working enzymes. The results were revealing. When the computer models looked at the slightly damaged enzymes, they failed almost completely. Even though the catalytic machinery had been destroyed, the models confidently assigned the original, correct job description to the broken enzymes. In fact, for the least damaged decoys, the models made false predictions more than 90 percent of the time. This suggests that the models are not actually learning the physical rules of how an enzyme works; instead, they are relying on a shortcut. They are looking at the overall similarity to known enzymes and assuming that if it looks like a duck, it must be a duck, even if the duck has been broken.
Even the most advanced model, which had been trained to recognize broken enzymes, struggled to identify the specific damage when it was focused on the active site. While this model did perform better when the damage was more severe, it still missed the targeted disruptions that rendered the enzyme useless. The study demonstrates that current computer programs are highly vulnerable to these phylogenetic shortcuts, meaning they are easily fooled by the family history of a protein rather than its actual function. The researchers conclude that to build truly reliable tools for understanding biology, the next generation of models must be trained with examples that include these structure-aware, broken versions. Only by teaching the computers to recognize the difference between a functional machine and a broken one, rather than just matching patterns, can we develop models that are robust enough to handle the complex reality of enzyme function.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.