← Latest papers
🧬 biology

Entropy-Driven Alignment-Free Phylogenetic Analysis via K-mer Encoding, DNA Physicochemical Properties, and Rough Set Feature Selection

This study proposes an alignment-free phylogenetic analysis framework that integrates k-mer statistics with DNA physicochemical properties and strand symmetry into a feature vector, which is then optimized using an enhanced rough set-based selection method to achieve efficient and interpretable whole-genome evolutionary inference.

Original authors: Yong Liu, Tongliang Xia, Xu Yan, He Wang, Xinyang Jiang, Shujie Gao, Shusheng Li, Yan Yang

Published 2026-08-20
📖 4 min read☕ Coffee break read

Original authors: Yong Liu, Tongliang Xia, Xu Yan, He Wang, Xinyang Jiang, Shujie Gao, Shusheng Li, Yan Yang

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

To understand the history of life, scientists often look at the molecular instructions written inside every living cell: the DNA sequence. For decades, the standard way to compare these instructions between different species has been to line them up side by side, letter by letter, much like aligning two long paragraphs of text to see where they match and where they differ. This method, known as multiple sequence alignment, works well when the sequences are short and very similar. However, as technology has advanced, researchers can now read entire genomes, which are millions of letters long. When these massive sequences are compared, especially among organisms that have shuffled their genetic material or evolved rapidly, the traditional alignment method becomes slow, computationally heavy, and sometimes impossible to perform accurately. This has led scientists to seek "alignment-free" methods, which try to find evolutionary relationships without forcing the sequences into a rigid line-up. Instead of matching every letter, these newer approaches look at the overall patterns and frequencies of small chunks of DNA to infer how closely related two organisms are.

In a recent study, a team of researchers from China proposed a new way to refine these alignment-free techniques to make them both faster and more accurate. They recognized that while many existing methods simply count how often certain short DNA patterns appear, they often miss two critical pieces of information: the specific chemical properties of the DNA building blocks and the order in which those patterns appear. To solve this, the researchers developed a system that translates a DNA sequence into a set of seven different "maps." Each map highlights a different physical characteristic of the DNA, such as whether a letter is part of a strong or weak chemical bond, or whether it belongs to a specific chemical family. By converting the original DNA sequence into these seven distinct binary patterns, they could then count the frequency of short word-like segments within each map. This process created a detailed, multi-layered profile of the genome that preserved both the chemical nature of the DNA and the specific order of its parts.

However, this detailed profile created a massive amount of data, containing thousands of potential patterns, many of which were redundant or unhelpful for determining evolutionary history. To cut through this noise, the team applied a mathematical strategy known as rough set theory. Think of this process as a highly efficient filter that sorts through a large pile of clues to find only the ones that actually help distinguish between different groups. The researchers tested this filtering system using a training set of eleven hantavirus strains, a type of virus found in East Asia. By analyzing how well different patterns could separate the viruses into their known types, the system identified a specific set of 3,005 patterns out of the original 14,336 possibilities. These selected patterns became the core "fingerprint" used to analyze the viruses.

The team then put their method to the test on three completely different groups of viruses to see if it could correctly reconstruct their family trees. First, they analyzed thirty-three coronavirus genomes, including the strains responsible for SARS and the more recent SARS-CoV-2. The method successfully grouped the viruses into four distinct branches that matched the known scientific classification, clearly separating the SARS-related viruses from the others and even distinguishing between the original SARS virus and the newer SARS-CoV-2. Next, they examined thirty-nine strains of the Japanese encephalitis virus. Again, the method correctly sorted the viruses into their four known genetic groups, mirroring the results obtained by traditional, much slower alignment methods. Finally, they analyzed fifty-five strains of the H7N9 avian influenza virus. The resulting tree showed two major lineages, one from North America and one from Eurasia, and correctly grouped viruses that were isolated in the same geographic regions and time periods.

Throughout these tests, the new approach proved to be remarkably fast. While other graphical methods for analyzing long DNA sequences can take several minutes, this method processed the longest sequences in the study, which were over 31,000 letters long, in just four seconds. The researchers noted that their approach does not require the complex and time-consuming step of aligning sequences, yet it still captures the essential biological signals needed to build accurate family trees. While the method is powerful, the authors acknowledge that it does not provide a visual picture of the differences between sequences in the way some other tools do, and it does not explicitly model complex genetic shuffling events. Nevertheless, the study demonstrates that by combining chemical properties with smart data filtering, scientists can create a tool that is both efficient and reliable for tracing the evolutionary history of viruses and other organisms across entire genomes.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →