← Latest papers
💻 bioinformatics

nf-cavalier: A Nextflow Pipeline for Rare Disease Variant Prioritization and Reporting

nf-cavalier is an open-source Nextflow pipeline that automates the annotation, filtering, and visual reporting of genomic variants to facilitate the identification of causal mutations in rare Mendelian diseases.

Original authors: Munro, J. E., Reid, J., Bahlo, M. E., Bennett, M. F.

Published 2026-08-10
📖 3 min read☕ Coffee break read

Original authors: Munro, J. E., Reid, J., Bahlo, M. E., Bennett, M. F.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are a detective trying to solve a mystery, but instead of looking for a missing person, you are hunting for a tiny typo in a massive instruction manual. This manual is your DNA, the code that tells your body how to build and run itself. Sometimes, a single letter in this code gets swapped, deleted, or added, and that tiny mistake can cause a rare disease. The problem is that your DNA manual is enormous—about 3 billion letters long—and every time we read it, we find millions of differences between people. Most of these differences are harmless, like a typo in a recipe that doesn't change the taste of the cake. But finding the one specific typo that actually causes the illness is like finding a needle in a haystack the size of a mountain. Scientists have to use powerful computers to sort through millions of these "typos," check if they are common in healthy people, see if they run in families, and predict if they are dangerous. It's a lot of work, and doing it manually is slow and prone to errors.

This is where a new tool called nf-cavalier comes in. Think of it as a super-smart, automated robot assistant for genetic detectives. The paper introduces this tool as a pipeline—a step-by-step assembly line—that takes the raw data from a patient's DNA scan and automatically filters out the millions of harmless typos to find the few "suspects" that might be causing the disease. It doesn't just guess; it uses a set of customizable rules to check if a typo is rare, if it breaks an important gene, and if it fits the family's inheritance pattern (like whether it came from one parent or both). Once it narrows down the list, it doesn't just spit out a boring spreadsheet. Instead, it creates a colorful, interactive report—like a digital detective board with clickable links and even PowerPoint slides—that lets doctors and researchers quickly review the evidence and decide what to do next.

The authors built this tool because existing methods were often fragmented, requiring researchers to juggle different software programs that didn't talk to each other, or they relied on expensive, complex servers that many labs couldn't afford. nf-cavalier is designed to be flexible and free, running on standard computers or powerful supercomputers without needing a permanent internet connection or a database. It handles all types of genetic errors, from tiny single-letter swaps to large chunks of DNA that are missing or duplicated. The paper shows that this tool has already been used successfully in real-world studies to diagnose patients with rare diseases, including those with epilepsy and developmental delays. In one study involving 242 people, it helped find the cause for 37 of them. In another, it helped solve a mystery for 16 out of 50 families. The tool doesn't automatically declare a diagnosis; instead, it acts as a highly efficient filter that highlights the most promising clues, making the job of human experts faster, more consistent, and easier to share with colleagues. By automating the heavy lifting of data sorting and presenting the results in a user-friendly format, nf-cavalier helps teams focus on the most important part of the job: understanding the disease and helping the patient.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →