OpenAntigens: a structure-aware database for antigen construct design across the human cell-surface and secreted proteome
OpenAntigens is a free, no-login database that streamlines antigen construct design for the human secreted and cell-surface proteome by integrating diverse structural, topological, and comparative data into a unified, interactive workflow to generate reproducible and cross-reactivity-aware design suggestions.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
In the world of modern medicine and biology, scientists often need to study specific proteins to understand how diseases work or to develop new treatments. Many of these proteins sit on the surface of cells or float freely in the body's fluids, acting as signals, doors, or anchors. To study them in a lab, researchers must first create a physical copy, or a "construct," of just the right piece of that protein. This is a tricky task because proteins are long, complex chains that fold into specific three-dimensional shapes. If a scientist cuts the chain in the wrong place, they might end up with a floppy, useless fragment that doesn't hold its shape, or they might accidentally include a part that is buried inside the cell membrane, which cannot be studied in a standard test tube.
The challenge is that a single protein can be described in many different ways across various scientific databases. One source might list where the protein starts and ends, another might show a computer-generated guess of its shape, and a third might list which parts are known to interact with other molecules. Before a scientist can order the DNA needed to make their protein copy, they have to manually gather all these scattered clues, check if they agree, and decide exactly where to cut the chain. This process is slow, prone to human error, and often requires juggling dozens of different websites and files just to answer a single, basic question: which part of this protein should I build?
A new resource called OpenAntigens aims to solve this bottleneck by bringing all that information together into one place. Developed by a team at the Institute for Protein Innovation, this database focuses on 5,328 human proteins that are either secreted by cells or sit on the cell surface. Instead of leaving researchers to piece together clues from multiple sources, OpenAntigens automatically gathers data from major biological archives, including records of protein shapes, known family relationships, and disease associations. It then uses this information to generate a detailed report for each protein, suggesting exactly which sections would make the best candidates for laboratory study.
The system works by looking at the protein's structure and its biological context. For proteins that have a predicted 3D shape from computer models, the database analyzes the confidence of that model to find stable, folded regions. It distinguishes between parts of the protein that are likely to be rigid and useful, and parts that are floppy or disordered. Based on this analysis, it offers five different types of suggestions for each protein. Some suggestions cover the entire outer section of the protein, which is useful if a scientist wants to study how the whole shape interacts with other molecules. Others focus on specific, well-defined domains or regions that have been seen in previous experiments. The system also provides "strict" suggestions that break the protein down into its smallest, most stable building blocks, which can be helpful for creating very focused tools.
Beyond just suggesting where to cut, the database helps scientists think about safety and specificity. It checks for "liabilities," such as unpaired chemical groups that might cause the protein to stick to itself in unwanted ways, or sites where the cell naturally cuts or modifies the protein. Crucially, it also looks at how similar the human protein is to proteins found in mice and monkeys, which are common animals used in research. By showing exactly which parts of the protein are identical across species, the database helps scientists decide if a tool they build for humans will also work in animals, or if they need to avoid certain areas to prevent the tool from reacting to the wrong target.
The result is a massive collection of over 55,000 specific design suggestions and nearly 150,000 comparisons to other proteins, all organized into a free, easy-to-use website. For any given protein, a user can see a visual map of the suggested cuts, the confidence in the shape, and the sequence of the amino acids. If a compatible computer model exists, the website even includes an interactive tool that lets users drag their hands across the protein's shape to select their own boundaries, watching in real-time as the system warns them about potential chemical issues or missing parts. This interactive feature allows researchers to refine the suggestions to fit their specific needs without needing to write complex code or manually cross-reference dozens of documents.
While the database provides a powerful starting point, the authors are careful to note that it does not guarantee that a protein will be easy to make or that a tool built from it will work perfectly. The suggestions are based on existing data and computer predictions, not on physical experiments performed by the database itself. Some proteins, particularly those that span the cell membrane multiple times or those that require a partner protein to function correctly, remain difficult to study, and the database flags these cases so scientists know to proceed with extra caution. The tool is designed to handle the heavy lifting of information gathering and initial planning, freeing researchers to focus on the actual experiments.
By organizing thousands of proteins into clear, actionable reports, OpenAntigens turns a chaotic, manual process into a streamlined workflow. It allows scientists to move faster from the idea of studying a protein to the physical act of building it. Whether the goal is to develop a new antibody, create a diagnostic test, or understand the structure of a disease-related protein, this resource provides a single, reliable map to navigate the complex landscape of human proteins, ensuring that the first step in the journey is as informed and precise as possible.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.