Machine-Readable Database for Genotype-Phenotype Analyses in Rare Diseases Using FKTN-related Congenital Muscular Dystrophies as a Prototype
This paper presents a machine-readable database aggregating published literature on FKTN-related congenital muscular dystrophies, featuring harmonized genotype-phenotype data and a new clinical severity scale to facilitate genotype-phenotype correlation and serve as a pilot for similar rare disease analyses.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Technical Summary: Machine-Readable Database for Genotype-Phenotype Analyses in FKTN-related Congenital Muscular Muscular Dystrophies
Problem Statement
FKTN-related congenital muscular dystrophies (CMDs) represent a highly heterogeneous, ultra-rare subset of -dystroglycanopathies. While the FKTN gene is a well-defined biallelic cause of these disorders, the clinical literature is fragmented across nearly 100 manuscripts describing disparate populations and varied phenotypes. Currently, this data exists primarily in unstructured text within case reports and series, lacking harmonization. This absence of a comprehensive, machine-readable database hinders the ability to predict genetic pathogenicity, understand pathogenic mechanisms, and explore therapeutic opportunities for these rare conditions. Existing genomic databases (e.g., ClinVar, LOVD) often lack detailed, standardized phenotypic information linked to specific genotypes.
Methodology
The authors employed a rigorous, multi-stage pipeline to extract, harmonize, and structure data from the literature:
Literature Identification and Extraction:
- An iterative search of PubMed (January 2022 – April 2024) using terms related to FKTN, fukutin, dystroglycanopathy, and muscular dystrophy identified 288 manuscripts.
- After filtering for relevance and inclusion criteria (patients with FKTN variants and clinical data), 92 manuscripts were retained, covering 749 patients (386 individual case reports/series and 363 grouped patients).
- Data was manually extracted into a spreadsheet ("catalogue") organized by patient rather than manuscript to prevent duplication. Extracted fields included variants, sex, ethnicity, age of onset, clinical manifestations across four organ systems (muscle, neurological, cardiac, ophthalmological), and treatment outcomes.
Genotype Harmonization:
- Variants were standardized to the Genome Reference Consortium Human Build 38 (GRCh38/HG38) format.
- Due to the existence of 10 distinct FKTN transcript isoforms in NCBI and the frequent lack of isoform specification in older literature, variants were cross-referenced against all ten isoforms using tools including VariantValidator, Varsome, Mutalyzer 3, and NCBI Nuccore.
- Conflicts were resolved by prioritizing clinically validated databases (LOVD, ClinVar) or reconstructing original locations from paper specifics.
- Data was converted to a standard Variant Call Format (VCF), with specific modifications for insertions and deletions. Haplotype notations that could not be mapped to specific variants were excluded.
Phenotype Harmonization (NLP Pipeline):
- Natural language clinical descriptions were converted into Human Phenotype Ontology (HPO) terms using PhenoTagger.
- To address the tool's inability to detect negations (e.g., "no muscle weakness" being misidentified as a positive phenotype), the authors integrated NegBio and developed custom scripts. These scripts utilized a dictionary of negation terms to filter out false positives.
- This combined approach filtered over 90% of non-pertinent positive information, reducing the initial 505 unique HPO terms to 341 relevant phenotypes after manual curation to remove confounding variables (e.g., steroid-induced symptoms).
Clinical Severity Scale Development:
- A new clinical severity scale was developed based on the maximal motor milestone achieved in a patient's life, adapting the modified Ueda classification.
- Scale Definitions:
- 1 (Mild): Independent ambulation achieved.
- 2 (Moderate): Independent sitting or standing achieved.
- 3 (Severe): Never achieved independent sitting, standing, or head control.
- C: Too young to determine milestone.
- X: Insufficient information.
- Binary indicators were also assigned for the presence/absence of neurological, cardiac, and ophthalmological pathology.
Key Results
- Dataset Composition: The final machine-readable dataset includes 749 patients from 92 manuscripts. After excluding "patients in ranges" (PIR) to ensure individual data integrity, the analysis focused on specific patient records.
- Genetic Findings: 52 unique genotype combinations were identified. The most prevalent was homozygosity for the Japanese ancestral founder variant (3-kb retrotransposon insertion in the 3' UTR). Other notable variants included the Ashkenazi founder variant (c.1167dup) and variants in Turkish populations.
- Phenotypic Distribution:
- Severity: 60 patients were classified as mild (1), 157 as moderate (2), and 60 as severe (3). 11 were too young (C), and others lacked sufficient data (X).
- Organ Systems: Muscle-related symptoms affected 337 patients, neurological (CNS) symptoms affected 173, ophthalmological symptoms affected 57, and cardiac symptoms affected 51.
- NLP Performance: The integration of NegBio and custom scripts successfully filtered over 90% of non-pertinent positive information during phenotype extraction, significantly improving data accuracy compared to using NegBio alone (~60% filter rate).
Significance and Claims
The paper presents a comprehensive literature extraction and harmonization approach for FKTN-related CMDs, serving as a pilot for similar data extraction in other rare monogenic diseases.
- Methodological Blueprint: The authors position this work as a "blueprint" for data curation in other rare monogenic diseases. The semi-automated pipeline (manual extraction + NLP harmonization + custom filtering) is designed to be replicable for other disorders with similar literature volumes (~200 papers).
- FAIR Data Principles: The dataset transforms unstructured historical literature into FAIR (Findable, Accessible, Interoperable, Reusable) data, allowing integration into broader rare disease registries.
- Clinical Utility: The newly developed clinical severity scale provides a standardized metric to categorize phenotypic data, facilitating genotype/phenotype correlation studies.
- Future Automation: The authors modestly note that while current automation is limited by the need for manual tagging and complex variant harmonization, this dataset serves as a control and training ground for future Large Language Models (LLMs) to achieve fully autonomous literature extraction and curation.
The work is supported by the Intramural Research Program of the NIH and is made available via the RARe-SOURCE® database and associated GitHub repositories.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.