← Latest papers
📄 bioengineering

ProtGPT3: an Open-source family of Promptable and Aligned Protein Language Models

This paper introduces ProtGPT3, an open-source family of promptable and aligned protein language models that demonstrates few-shot prompting and inference-time steering as scalable, effective alternatives to fine-tuning for generating functional and diverse protein sequences.

Original authors: Garibbo, M., Boxo Corominas, G., Stocco, F., Illanes Vicioso, R., Middendorf, L., Ferruz, N.

Published 2026-06-08
📖 3 min read☕ Coffee break read

Original authors: Garibbo, M., Boxo Corominas, G., Stocco, F., Illanes Vicioso, R., Middendorf, L., Ferruz, N.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you have a giant, magical cookbook that knows every recipe for every dish in the universe. In the world of biology, this "cookbook" is a computer program that understands the "recipes" for building proteins—the tiny machines that keep living things alive.

The paper introduces a new, open-source version of this cookbook called ProtGPT3. Here is how it works, explained through simple analogies:

1. The Problem: Too Much Chaos

Previously, these computer programs were like a chef who could invent millions of new dishes, but they often made up nonsense meals that no one could eat. Scientists wanted to guide the chef to make specific types of dishes (like "spicy pasta" or "vegan desserts"), but it was hard to give clear instructions without having to retrain the chef from scratch every time.

2. The Solution: A Flexible Family of Chefs

The authors created a whole family of these protein chefs, ranging from small, quick ones to massive, super-smart ones (up to 10 billion "brain cells"). They made them open for everyone to use, just like a public library of recipes.

3. Two New Ways to Give Instructions

The paper highlights two clever ways to tell these chefs what to cook, without needing to retrain them:

  • The "Prompt" Method (Few-Shot Prompting): Instead of teaching the chef a whole new language, you just show them a few examples of the dish you want. It's like handing a chef a photo of a "perfect chocolate cake" and saying, "Make something like this." The paper found this method is just as good as the old, slow way of retraining the chef, but much faster and easier.
  • The "Alignment" Method (Fine-Tuning): This is like teaching the chef a set of strict rules about what makes a "good" dish. The researchers taught the models to avoid making "bland" or repetitive recipes (low-complexity) while still keeping the flavors diverse. This ensures the generated proteins are high-quality and stable.

4. The "MSA" Superpower

Some of these chefs have a special ability called MSA (Multiple Sequence Alignment). Think of this as a chef who doesn't just look at one recipe, but can look at a whole stack of similar recipes from different cultures to understand the essence of a dish.

  • The Result: When the team tried to design a specific protein to remove fluorine (a "low-data defluorinase"), the chef using this "stack of recipes" method (ProtGPT3-MSA) did a better job than the chefs who were retrained from scratch.
  • Real-World Proof: They didn't just simulate this on a computer; they actually built the protein in a lab. It worked: the protein dissolved properly and was successfully expressed, proving the computer design was real and functional.

5. Steering the Ship in Real-Time

Finally, the paper describes a new trick called "Feynman–Kac inference." Imagine you are driving a car toward a destination. Usually, you pick a route and stick to it. This new method is like having a GPS that constantly adjusts your steering wheel while you are driving to make sure you stay exactly on the path to your target, even if the road gets tricky. This allows scientists to steer the protein generation toward a specific goal right at the moment of creation.

In short: The authors built a versatile, open-source toolkit that helps scientists design new proteins more easily. They showed that you don't always need to retrain the AI; sometimes, just giving it good examples or using a "recipe stack" approach works better, faster, and produces results that actually work in the real world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →