Joint optimisation of amino acid and coding sequence for de novo designed proteins
This paper introduces JANUS, a joint optimization framework that simultaneously designs amino acid sequences and coding sequences for de novo proteins by leveraging inverse-folding entropy to eliminate experimental liabilities and improve synthesis success rates, outperforming traditional methods that treat these tasks separately.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
In the quiet, microscopic world of life, the instructions for building a protein are written in a double language. The first language is the protein itself, a chain of amino acids that folds into a specific shape to perform a job, like a key fitting a lock or an enzyme speeding up a chemical reaction. The second language is the gene, a strand of DNA that acts as the blueprint for that protein. For millions of years, nature has written these blueprints, and every natural protein comes with a gene that has been tested and refined by evolution. But in modern science, researchers are now designing entirely new proteins from scratch, creating shapes that nature never made. The challenge is that while scientists have become very good at imagining these new shapes, they often struggle to build the genes needed to make them. A gene is not just a passive copy of the protein; it is a complex instruction set where the specific choice of letters matters for how well the gene is read, how stable the message is, and whether the cell can actually produce the protein without choking on it.
For a natural protein, the gene is a known quantity, a reliable partner that has survived the test of time. For a brand-new, designed protein, there is no such partner. Scientists must invent the gene from nothing. Until now, the standard approach has been to treat these two tasks as separate steps. First, a computer designs the perfect protein shape. Then, a different computer program takes that fixed shape and tries to find the best possible gene to build it, assuming the protein sequence cannot be changed. This paper argues that this separation is a mistake. The researchers show that the protein sequence is not a fixed target but a flexible one, and that by allowing the protein sequence to shift slightly while designing the gene, they can solve problems that a standard gene designer cannot touch. They developed a new method that designs the protein and its gene together, finding a solution that is better for the cell, easier to manufacture, and more likely to work in the lab.
The researchers began by looking at a massive collection of 862 protein backbones, including hundreds of new designs created by other scientists. They asked a simple question: how much freedom do we actually have when we design a gene for a specific shape? They found that for almost every position in these new proteins, there are several different amino acids that would fit the shape just as well. In fact, the computer models showed that for a typical design, there is a vast amount of hidden flexibility, a kind of "wiggle room" where the protein could be built with different building blocks without losing its shape. The standard method throws this flexibility away, picking one sequence and locking it in before the gene is even considered. The new approach, which the authors call JANUS, keeps that flexibility open. It treats the protein and the gene as a single, connected puzzle, searching for the best combination of both at the same time.
When they tested this joint approach on 447 published protein designs, the results were striking. They discovered that 99.1% of these designs carried at least one hidden flaw, a "liability" that could cause the protein to fail when the cell tries to make it. These flaws included sequences that were too repetitive, regions that were too sticky and prone to clumping together, or signals that told the cell to destroy the protein. A standard gene optimizer, which can only change the gene letters while keeping the protein fixed, could not fix these problems. It was like trying to fix a leaky roof by only rearranging the furniture inside; the structure itself was the issue. By allowing the protein sequence to change slightly, the new method could remove these flaws. For three of the five types of flaws they looked at, the cost of fixing them was incredibly small, requiring only a tiny amount of change to the overall design.
The most surprising and practical finding was not about the biology of the protein, but about the logistics of making it. Commercial companies that synthesize genes for scientists have strict rules about what they will build. They cannot make genes with certain patterns of letters because those patterns are difficult to manufacture or prone to errors. The researchers found that when they used the old method of designing the protein first and then the gene, the resulting genes violated these manufacturing rules more than half the time. When they used the new joint method, the number of violations dropped by more than half. This means that a gene that a vendor would previously refuse to make could now be ordered and built. This is a critical bottleneck; if a gene cannot be ordered, the project stops before the biology is even tested. The new method essentially clears the path for these new proteins to leave the computer and enter the real world.
The study also looked at how much the protein sequence actually needed to change to get these benefits. They found that for the vast majority of designs, the method only needed to swap out a few amino acids to achieve a much better result. In fact, they showed that a very small amount of flexibility was enough to capture almost all the possible improvement. This suggests that the "perfect" protein sequence is not a single, rigid point, but a landscape with many good options. The joint method simply finds the one spot on that landscape where the protein is stable, the gene is easy to read, and the manufacturing rules are satisfied. They also tested whether this approach worked better in different types of cells, such as bacteria versus human cells, and found that the method adapted successfully to both, though the specific changes required were different for each host.
Despite these successes, the researchers were careful to note where their method did not provide a massive advantage. For one specific part of the gene that controls how fast the protein production starts, they found that most of the improvement could be achieved just by changing the gene letters without touching the protein at all. This confirms that for some tasks, the old way of doing things is still quite good. However, for the broader set of problems that involve the protein's stability and its interaction with the cell's cleanup systems, the joint approach is essential. The study concludes that the field of protein design has been treating the gene as a secondary afterthought, but for de novo proteins, the gene is a primary constraint. By solving for the protein and the gene together, scientists can create designs that are not just theoretically sound, but practically viable, turning more of their digital creations into real, working molecules.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.