Detecting and typing Chlamydia trachomatis strains in metagenomes using the MetaChlam pipeline
The authors developed MetaChlam, an automated Nextflow pipeline that integrates multiple bioinformatics tools to accurately detect and type *Chlamydia trachomatis* strains in metagenomic samples with as few as 250 reads, overcoming limitations of traditional genotyping methods and revealing unexpected contamination in non-infectious datasets.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Bacteria that live inside human cells present a unique challenge for scientists trying to track them. Chlamydia trachomatis is one such organism, a microscopic invader that cannot survive outside its human host. It is a leading cause of sexually transmitted infections and a major driver of a blinding eye disease known as trachoma. While the bacteria are all the same species, they are not all identical; different strains prefer to infect different parts of the body, such as the cervix, the rectum, or the eye, and these preferences lead to different health outcomes. To understand which strain is present in a person, researchers usually need to isolate the bacteria and read its genetic code in detail. However, when scientists look at samples taken directly from the body, the amount of bacterial genetic material is often so small that it gets lost in the vast sea of human DNA and other microbes. This makes it difficult to identify the specific strain causing an infection using standard methods.
To solve this problem, a team of researchers developed a new digital tool called MetaChlam. This software is designed to sift through complex genetic samples, known as metagenomes, to find and identify Chlamydia trachomatis even when the bacteria are present in very low numbers. The researchers tested their system using a collection of 109 known bacterial genomes to determine exactly how similar two strains need to be before they can be told apart. They found that a specific threshold of genetic similarity could reliably distinguish one strain from another. By combining four existing software programs with their own custom databases, they created an automated pipeline that can process data without constant human intervention. When they tested this pipeline on computer-generated samples that mimicked real-world conditions, it successfully identified the correct strains, whether the sample contained just one type of bacteria or a mixture of several.
The tool proved its worth when applied to real data from public scientific databases. In these tests, MetaChlam was more accurate than existing software at confirming that the bacterial reads it found were truly from Chlamydia trachomatis and not from something else. This precision is vital because the researchers made a surprising discovery during their analysis: genetic fragments from this strictly human-infecting bacteria appeared in samples from environments where the organism should not exist. These findings suggest that the bacteria, or at least its genetic traces, can end up in samples as contaminants, a detail that could confuse diagnosis if not properly accounted for. The study concludes that this new pipeline offers a reliable way to characterize and type these bacterial strains directly from complex genetic data, providing a clearer picture of the infections they cause. The software is now available for other scientists to use, helping to improve how these pathogens are identified in the future.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.