FIND: a software tool for identifying population-enriched pathogenic variants in gnomAD
The authors present FIND, a freely available web tool that identifies population-enriched pathogenic variants in the gnomAD database by detecting alleles with significantly higher frequencies in specific ancestry groups, thereby facilitating the discovery of founder mutations and population-specific disease burdens, including in historically underrepresented populations.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine you have a massive library containing the genetic "storybooks" of millions of people from around the world. This library is called gnomAD. Inside these books, there are tiny typos (genetic variants). Most typos are harmless, but some are dangerous "plot twists" that cause disease.
The problem is that some of these dangerous typos aren't scattered randomly. Instead, they are like family heirlooms that got passed down through specific groups of people who lived in the same village or shared a common ancestor long ago. Because these groups stayed somewhat isolated (a "bottleneck" and "endogamy"), these specific typos became much more common in that one group than anywhere else.
For a long time, finding these "heirloom typos" was like trying to find a specific needle in a haystack while wearing blinders. You might know the needle exists, but you couldn't easily see that it was ten times more common in one pile of hay than in all the others.
Enter FIND:
The authors built a digital magnifying glass called FIND (Founder candidates hidden IN Data). Think of it as a super-smart librarian who scans the entire library with a very specific rule:
- Look for the "bad" typos: It only checks the pages known to be dangerous (pathogenic or likely to break the story).
- Spot the imbalance: It asks, "Is this typo showing up in one specific group of people at least ten times more often than in every other group?"
- Ignore the noise: If a group is too small to be sure (fewer than five copies of the typo), the tool ignores them to avoid false alarms.
What did FIND find?
The team tested their new tool on five famous genes (like BRCA1 and BRCA2, which are well-known for cancer risks). The results were like a treasure hunt:
- The Knowns: It immediately found 12 famous "heirloom typos" that scientists already knew about, proving the tool works.
- The Recurring Suspects: It found 5 typos that were known to happen often in certain groups but had never been compared across the whole world before.
- The New Discoveries: It spotted 3 typos that no one had ever realized were concentrated in specific populations.
Why does this matter for everyone?
The tool didn't just look at the usual suspects. It also checked groups that have often been left out of genetic studies, like African American and admixed American populations. By cross-checking with another massive database called "All of Us," they confirmed that FIND can spot these hidden patterns in groups that have historically been ignored.
The Bottom Line:
FIND is a free, open-source tool (like a public utility) that helps scientists see the "hidden in plain sight" genetic risks that are specific to certain communities. It doesn't invent new medicine, but it shines a light on the existing data so we can finally see which dangerous typos are clustered in which families, making it easier to screen for them.
You can try the tool yourself on a website or look at the code on GitHub, just like a public park that anyone can visit.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.