← Latest papers
💻 bioinformatics

PG-LLM: Benchmarking General-Purpose Language Models for Protein Variant Ranking

The PG-LLM benchmark demonstrates that general-purpose language models, despite lacking access to structural data or multiple-sequence alignments, can effectively rank protein variants and outperform many established sequence-based predictors, though they still lag behind specialized biomolecular tools.

Original authors: Arora, R. K., Chen, L. T., Du, M., Marks, D., Church, G.

Published 2026-08-03
📖 4 min read☕ Coffee break read

Original authors: Arora, R. K., Chen, L. T., Du, M., Marks, D., Church, G.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are a master chef trying to invent a new recipe. You have a basic dish (a protein) that works well, but you want to tweak the ingredients (the amino acids) to make it taste even better, last longer, or work in a different kitchen. In the world of biology, this is called protein engineering. For years, scientists have used specialized computer programs—like a team of expert food critics who have memorized every recipe ever written—to guess which tweaks will work. These experts look at the family history of the dish (evolutionary data) and the physical shape of the pot (protein structure) to make their predictions.

Recently, a new type of "chef" has entered the kitchen: General-Purpose Language Models (LLMs). These are the same AI brains that can write poems, solve math problems, and chat about history. They are incredibly smart at understanding language and patterns. The big question scientists have been asking is: Can these general AI chefs, who haven't been specifically trained to be food critics, look at a recipe and a list of ingredient swaps, and tell us which new version will be the best? They don't have the family history books or the 3D models of the pots; they just have the text of the recipe and their own vast knowledge of how words (and in this case, amino acids) fit together.

This paper, titled "PG-LLM," sets up a massive taste-test to find out. The researchers created a benchmark called PG-LLM, which is like a giant cooking competition. They took 217 different "recipes" (proteins) and asked thirteen different AI chefs to rank 50 different versions of each recipe from "best" to "worst" based on how well they would work in a real experiment. The catch? The AI chefs weren't allowed to use any special protein tools, evolutionary maps, or 3D structures. They had to rely on their general smarts alone.

The results are a mix of "wow" and "not quite there yet." The top AI chef, Claude Opus 5, managed to get a score of 0.406 (on a scale where 1.0 is perfect). This is a pretty good score! It means the AI is actually quite good at guessing which protein tweaks will work, beating out 49 other specialized computer programs that have been around for years. In fact, it did better than almost all the "sequence-only" experts (the ones that just look at the list of ingredients without 3D models).

However, the AI chefs didn't win the whole competition. The very best specialized protein experts, like a program called VenusREM, still scored higher at 0.523. It's like the general AI is a very talented home cook who can guess the winner of a baking contest, but the professional pastry chef with decades of specific training still knows the secrets better.

The paper also discovered some interesting quirks about how these AI chefs think. When the researchers made the AI think harder and longer (giving it more "test-time compute"), the scores got better, but they eventually hit a wall. The AI also struggled more when the list of 50 recipes got too long to juggle in its mind, whereas the specialized experts didn't care how long the list was. Interestingly, the AI's performance didn't change much whether the protein had a deep family history or a shallow one, suggesting the AI is using a different kind of logic than the evolutionary experts.

Finally, to make sure the AI wasn't just relying on memorized answers from the internet, the researchers tested them on brand-new recipes that were published after the AI's training data cut-off. The results were similar: the AI could still reason about the new recipes, proving it's actually learning the rules of the game, not just reciting old answers.

In short, this paper shows that general-purpose AI is becoming a surprisingly capable "biological reasoner." It can already outperform many established tools and is getting closer to the top specialists. But for now, it hasn't quite replaced the need for the specialized, protein-focused experts. It's a powerful new tool in the lab, but the master chefs are still in the game.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →