← Latest papers
💻 bioinformatics

vep-rs: high-throughput Rust variant annotation with population-scale concordance to Ensembl VEP

vep-rs is a high-throughput Rust reimplementation of Ensembl VEP that achieves near-perfect concordance with the original tool across millions of variants while delivering 78x to 284x faster performance and correcting several documented defects in the Perl implementation.

Original authors: Porter, M., Borkowski, R.

Published 2026-09-28
📖 5 min read🧠 Deep dive

Original authors: Porter, M., Borkowski, R.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

In the modern era of medicine, reading a person's genetic code has become fast and affordable, but understanding what that code means remains a slow and expensive bottleneck. When scientists sequence DNA, they first identify tiny differences, or variations, between an individual's genome and a standard reference. These variations are the raw data. The next, crucial step is to interpret them: to determine if a specific change disrupts a gene, alters a protein, or is harmless. This interpretation relies on a massive, complex set of rules that maps every possible genetic change to a biological consequence. For over a decade, the scientific community has relied on a single, widely used software tool called Ensembl VEP to perform this mapping. It is the standard dictionary for geneticists, translating raw genetic data into a language doctors and researchers can use to make sense of disease. However, this standard tool was built with an older programming language that struggles with speed. As the volume of genetic data has exploded, the time it takes to run this tool has become a major hurdle, slowing down everything from research projects to clinical diagnoses.

A team of researchers at Natera, Inc., has now built a new tool called vep-rs to solve this speed problem without sacrificing accuracy. They rewrote the entire logic of the standard tool using a modern programming language known for its efficiency and safety. To ensure their new tool was a perfect replacement, they did not just guess; they tested it against the original software using a massive dataset containing over 260 million genetic variations. This is a scale of testing that had never been done before for this specific task. The results were striking: on the vast majority of genetic changes, the new tool produced the exact same answers as the original, but it did so between 78 and 284 times faster. In practical terms, a task that might have taken the original tool an hour to complete was finished by the new tool in less than a minute. This leap in speed means that laboratories can now process population-scale genetic data almost instantly, removing a critical barrier to faster medical insights.

The researchers were careful to verify that their new tool was not just fast, but also correct. They compared the output of their software against the original tool across six different large datasets, including data from the ClinVar database and the gnomAD project. On the most common types of genetic changes, known as single-letter swaps and small insertions or deletions, the new tool matched the original tool's results in more than 99.99 percent of cases. The few times the two tools disagreed, the researchers investigated each instance closely. They found that in most of these rare disagreements, the original tool was actually making a mistake, such as contradicting itself or producing results that depended on how the data was grouped together. The new tool, by contrast, remained consistent and logical. In fact, once the known errors of the original tool were set aside, the new tool matched the original's intended logic perfectly across all tested datasets.

The study also looked at more complex genetic changes, such as large structural variations where chunks of DNA are deleted or rearranged. Here, the new tool still performed exceptionally well, matching the original tool's results with high precision, though the complexity of these variations meant the agreement was slightly lower than for simple changes. Even in these difficult cases, the new tool was significantly faster. The researchers also compared their tool to another existing fast alternative, but found that the other tool missed a large number of results and failed to match the standard tool's vocabulary, making it unsuitable for pipelines that rely on the established dictionary of genetic terms. The new tool, vep-rs, was designed to speak the exact same language as the original, allowing scientists to switch to the faster version without having to rewrite their analysis software or change their filters.

Beyond raw speed and accuracy, the new tool offers a level of reliability that the original tool lacks. The researchers documented five specific types of errors where the original tool behaves inconsistently. For example, the original tool sometimes assigns a genetic change to a chromosome it does not actually belong to, or it produces different results depending on how many other changes are in the same file being processed. The new tool eliminates these issues entirely, providing a stable and predictable output. This consistency is vital for clinical settings, where a result must be the same every time it is run, regardless of the batch size or the order of the data. By removing these contradictions, the new tool not only speeds up the process but also improves the quality of the data itself.

The development of vep-rs represents a significant shift in how genetic data is handled. By proving that a complete rewrite of a critical scientific tool can achieve near-perfect agreement with the original while delivering massive gains in speed, the researchers have opened the door to analyzing genetic data at a scale that was previously impractical. The tool is now available for public use, allowing any laboratory to adopt it immediately. As the cost of sequencing continues to drop and the volume of data grows, the ability to process this information quickly and accurately becomes the limiting factor in medical progress. This new software removes that bottleneck, ensuring that the insights hidden in our genetic code can be discovered and acted upon with unprecedented speed.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →