AlphaFold Database expands to proteome-scale quaternary structures
This paper describes the expansion of the AlphaFold Protein Structure Database to include 1.81 million high-confidence, proteome-scale predictions of homo- and heterodimeric complexes across 4,777 organisms, revealing novel structural topologies and conserved ancient assemblies that significantly extend beyond existing experimental coverage to facilitate functional discovery.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine the human body (and every other living thing) as a massive, bustling city. Inside this city, the workers are proteins. For a long time, scientists thought of these workers as solo employees, each doing their job in isolation. We even built a giant digital library, called the AlphaFold Database, that contains the blueprints for millions of these solo workers, showing us exactly what they look like in 3D. But here's the catch: in the real city, proteins rarely work alone. They are social creatures that constantly team up, shake hands, and form complex machines to get things done. These teams are called "complexes." Until now, our library was missing the blueprints for these teams. We knew who the workers were, but we didn't know how they held hands or built their machines. This matters because if you want to fix a broken machine or understand how a city runs, you need to see the workers together, not just standing alone.
Now, a massive collaboration between scientists at the European Molecular Biology Laboratory, Seoul National University, NVIDIA, and Google DeepMind has finally opened the doors to this missing part of the library. They didn't just guess; they used supercomputers to predict the 3D shapes of over 31 million potential protein teams. Out of this mountain of predictions, they filtered out the shaky guesses and kept 1.81 million high-confidence models of protein pairs working together. Think of it as upgrading the library from a collection of solo portraits to a vast gallery of group photos. They found that while some teams look very similar to ones we've already seen under microscopes, many others are brand new, revealing structures that were completely invisible when we only looked at the proteins alone. This new resource is now free for everyone to explore, offering a fresh look at how life's molecular machinery is actually assembled.
The Big Leap: From Solo Acts to Team Sports
For years, the AlphaFold Protein Structure Database (AFDB) was like a photo album of individual actors. It gave us the 3D shapes of single proteins with incredible accuracy, revolutionizing how we understand biology. But biology is rarely a solo act. Proteins usually function by grabbing onto other proteins, forming pairs or larger groups to perform tasks like sending signals, building cell walls, or copying DNA. The problem was that while we had the photos of the actors, we didn't have the photos of the duets.
This new paper describes a massive expansion of that database. The team predicted the structures of 31 million candidate protein pairs (called dimers) from 4,777 different organisms, ranging from common lab bacteria to organisms critical for global health. They didn't just throw everything into the mix; they used a rigorous filtering process to find the most reliable predictions. After running these through a quality check, they identified 1.81 million high-confidence structures. These are the "gold standard" models where the computer is very sure the two proteins are actually holding hands in the way predicted.
Why Looking at Pairs Changes Everything
The most exciting part of this discovery is that seeing proteins as a team often changes the story entirely. Sometimes, a protein looks like a messy, unrecognizable blob when predicted alone, but when you put it next to its partner, it snaps into a perfect, clear shape.
The authors found several ways this happens:
- The "Puzzle Piece" Effect: Some proteins are like puzzle pieces that only make sense when they click together. For example, one protein from a slime mold (a type of microbe) looked like a broken, low-confidence mess on its own. But when modeled as a pair, the two halves swapped parts of their structure to complete each other's shapes, forming a perfect, high-confidence machine.
- The "Wall Builder" Effect: Some proteins are like bricks that need to be stacked to form a wall. A protein involved in cell cleanup (autophagy) looked like a wobbly, uncertain helix bundle on its own. But when paired with its twin, they formed a solid, coherent structure that clearly defined where the cell membrane should be, something the solo model couldn't figure out.
- The "Architect" Effect: Even when a solo protein looks good, adding a partner can fix how its different parts are arranged. One protein had a confident shape, but the distance between its two main sections was a guess. When modeled as a pair, the partner acted like a scaffold, locking the two sections into the correct position.
In short, the paper shows that for many proteins, the "real" shape only exists when they are in a complex. You can't understand the whole picture by looking at the parts in isolation.
What the Numbers Tell Us
The team didn't just make pretty pictures; they did the math to see how reliable these predictions are. They set strict rules for what counts as a "high-confidence" team. They found that:
- Homodimers (identical pairs): About 9.1% of the predicted identical pairs passed the high-confidence test. This is a lot! It suggests that when two of the same protein team up, the computer is quite good at figuring out how they fit.
- Heterodimers (different pairs): Only about 1.0% of the predicted different pairs passed the test. This is lower, and the authors suggest it's because they cast a wide net to find potential partners, meaning many of the pairs they tested might not actually interact in real life.
- Ancient Teams: When they grouped these structures by similarity, they found that the top 1% of the most common structures accounted for about 44% of all the high-confidence complexes. Even more fascinating, about 8.3% of these common structures were found across completely different branches of life (like bacteria, plants, and animals). This suggests that some of these protein teams are ancient, having been kept by evolution for billions of years because they work so well.
Beyond the Known World
One of the biggest takeaways is how much of this new world is new. When they compared their 1.81 million high-confidence models to the existing library of experimentally determined structures (the PDB), they found that 31.3% of the individual predictions didn't match anything we've ever seen under a microscope. At the level of groups (clusters), 63.6% of the unique structures had no known experimental match.
This means the database isn't just filling in the blanks; it's discovering entirely new territories. While many of these predictions are anchored in what we already know (giving scientists confidence to use them), a huge chunk extends into uncharted territory, offering hypotheses for structures that experimentalists haven't found yet.
The Bottom Line
This paper is a massive step forward, turning the AlphaFold Database from a collection of solo portraits into a dynamic gallery of molecular teamwork. By providing 1.81 million high-confidence models of protein pairs, the authors have given biologists a new tool to generate hypotheses about how proteins interact, how diseases might disrupt these teams, and how to design new drugs. They aren't claiming to have solved every protein interaction in the universe, but they have provided a scalable, reliable framework that bridges the gap between what we know and what we need to discover. The library is now open, and the group photos are ready for you to explore.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.