GeneResolver: A Traceable Hybrid Multi-Agent Pipeline for Etiology-Aware Gene Prioritization
GeneResolver is a traceable, privacy-oriented hybrid multi-agent pipeline that improves rare-disease diagnosis by first classifying case etiology to route evidence appropriately, thereby outperforming existing tools in both phenotype-only and combined phenotype-genomic benchmarks.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
For families navigating the labyrinth of a rare genetic disease, the path to a diagnosis is often a long and weary journey. Even with modern technology that can read a person's entire genetic code, finding the specific error responsible for an illness remains difficult. The challenge is not just in reading the code, but in understanding what it means. Doctors must sift through thousands of genetic variations, looking for the one that matches a patient's unique collection of symptoms. This process relies heavily on a precise description of those symptoms, often translated into a standardized list of medical terms. However, turning a doctor's messy, handwritten notes or a complex story of a patient's history into that clean, standardized list is a slow, manual task. Furthermore, many tools used to find the cause of a disease assume the problem comes from a single broken gene. This assumption can lead doctors down the wrong path if the illness is actually caused by a large chromosomal error, a mix of many small genetic factors, or something entirely outside the genes.
A new system called GeneResolver, developed by researchers in North Macedonia, attempts to solve these problems by acting as a smart, privacy-focused guide for doctors. Instead of jumping straight to ranking genes, this system first reads the patient's story to decide what kind of biological problem it is facing. It asks a simple but crucial question: is this likely caused by a single gene, or is it something else? If the answer is a single gene, the system proceeds to rank potential genetic culprits. If the answer is something else, it stops the gene search and instead offers evidence about the actual cause, such as a chromosomal issue or a non-genetic factor. The system is designed to work locally on a hospital's own computers, ensuring that sensitive patient stories and genetic data never leave the building, a critical feature for protecting privacy.
The researchers built this system as a team of specialized digital assistants working together. One assistant reads the raw clinical notes and carefully removes any personal details like names or addresses, replacing them with generic placeholders. Another assistant translates the remaining medical story into a structured list of symptoms using a standard medical vocabulary. A third assistant looks at genetic data if it is available, checking for errors in the genetic code. A central coordinator then takes all this information and makes a judgment call on the type of disease. If the system determines the case fits a single-gene pattern, it moves to a ranking phase where it weighs the symptoms against known genetic associations and the severity of any genetic errors found. If the system decides the case is not a single-gene disorder, it routes the information to a different branch that gathers evidence relevant to that specific type of problem, avoiding the mistake of forcing a single-gene explanation onto a complex case.
In testing this approach, the researchers found that the system could successfully distinguish between single-gene cases and other types of disorders with high accuracy. When asked to classify raw patient stories without any prior examples to learn from, the system correctly identified the type of disorder in 77 percent of cases. More importantly, it correctly separated single-gene cases from all other types 86 percent of the time. This ability to route cases correctly is vital because it prevents the system from wasting time or giving misleading results when a gene search is not the right tool. The system also performed well at the core task of finding the right gene when one was likely the cause. On a large set of nearly 4,700 cases where only symptom descriptions were available, the system placed the correct gene at the very top of its list more often than other widely used tools. When genetic data was added to the mix, the system's ability to find the correct gene improved even further, placing the right answer at the top in over half of the cases.
The study also highlighted the importance of keeping data private. The system uses a large language model to understand the context of medical stories, but this model runs entirely on local hardware within the hospital. This means the sensitive details of a patient's life and health are never sent to an external server. The researchers tested different methods for cleaning personal information from notes and found that a specific local tool was the most effective at removing identifying details while keeping the medical meaning intact. By combining this local privacy protection with a flexible workflow that adapts to the type of disease, GeneResolver offers a new way to approach rare diseases. It does not claim to replace the doctor; rather, it acts as a traceable, evidence-based partner that organizes complex information and suggests the most likely paths forward, leaving the final decision and interpretation to the human expert.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.