← Latest papers
🧬 biology

Information-theoretic profiling of structural organization in human genes

This study employs an information-theoretic clustering framework based on Kullback–Leibler cluster entropy to analyze the structural organization of specific human genes involved in epigenetic regulation and neurodevelopmental disorders, demonstrating that the resulting entropy profiles are robust and useful for comparative genomics and functional gene characterization.

Original authors: Chiara Panico, Filippo Gandino, Renato Ferrero, Anna Carbone

Published 2026-09-01
📖 5 min read🧠 Deep dive

Original authors: Chiara Panico, Filippo Gandino, Renato Ferrero, Anna Carbone

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

The human genome is a vast library containing the instructions for building and running a human being. It is written in a chemical code made of four letters, representing the building blocks of DNA. For decades, scientists have focused on reading the specific words in this book—the genes that code for proteins—to understand how life works. However, the spaces between these words, and the way the letters are arranged within them, hold a different kind of secret. These regions are not just empty filler; they contain complex patterns of organization that help control when and where genes turn on or off. To understand these hidden structures, researchers are turning to a branch of mathematics originally developed to measure information and uncertainty. By treating DNA sequences as data streams, they can look for statistical patterns that reveal how the genetic code is organized, without needing to align or compare the sequences to other species. This approach allows them to see the "texture" of the genome, identifying areas that behave in a highly ordered way versus those that appear more random.

In a recent study, researchers from the Politecnico di Torino in Italy applied this information-based approach to five specific human genes known to be involved in regulating how DNA is packaged and read. These genes, named ASH1L, NSD1, NSD2, NSD3, and SETD2, are crucial for development and are linked to various diseases when they malfunction. The team wanted to see if the mathematical "fingerprint" of these genes could reveal their internal structure. They started by converting the long strings of DNA letters into a series of numbers. To do this fairly, they grouped the DNA letters into two categories based on their chemical shape: those with two rings and those with one. This created a balanced numerical map of each gene, stretching from the start of the gene to a thousand letters before and after it, ensuring they captured the surrounding regulatory environment.

The researchers then analyzed these numerical maps by sliding a window along the sequence, looking at small segments at a time. For each segment, they measured how the letters clustered together and compared that pattern to a set of computer-generated models. These models represented different types of randomness, ranging from completely chaotic noise to patterns with long-range connections, similar to how a weather system might have correlations over time. The goal was to find which computer model best matched the real DNA sequence. If a segment of DNA behaved exactly like pure randomness, the match would be perfect. If it showed a specific, non-random structure, the mathematical distance between the real DNA and the random model would increase. This distance, or divergence, served as a signal of how structured that part of the gene was.

The results revealed a striking and consistent pattern across all five genes. The sections of DNA that code for proteins, known as exons, showed a strong, distinct signal. These coding regions consistently deviated from the random models, indicating they possess a specific, organized internal structure. In contrast, the non-coding sections, called introns, which make up the vast majority of the gene's length, behaved much more like the random, uncorrelated models. The mathematical measure effectively highlighted the difference between the "working" parts of the gene and the "spacer" parts. When the researchers compared these mathematical profiles to the known biological maps of the genes, the peaks in the signal lined up perfectly with the locations of the protein-coding exons. The valleys in the signal corresponded to the introns.

This finding suggests that the information-theoretic method can distinguish between coding and non-coding DNA without needing to know the biological function beforehand. The study also looked at other regulatory features, such as promoters and enhancers, which act as switches for the genes. The connection between the mathematical signal and these regulatory elements was less clear and varied from gene to gene, suggesting that while the coding structure is robustly captured by this method, the organization of regulatory switches is more complex and context-dependent. The researchers found that the strength of the signal in the coding regions was closely tied to the size of the gene and the amount of coding DNA it contained.

The study demonstrates that the structural organization of human genes leaves a detectable statistical signature. By using a method that measures how much a sequence differs from pure randomness, the researchers were able to map the internal architecture of complex genes. The approach proved robust, working consistently whether the data was analyzed in overlapping or non-overlapping sections. While the method clearly separates the coding from the non-coding regions, its ability to pinpoint specific regulatory switches is still developing. The work offers a new, complementary way to look at the genome, one that relies on the mathematical properties of the sequence itself rather than just comparing it to other known sequences. This could eventually help scientists identify functional regions in the genome more efficiently, providing a deeper understanding of how the genetic code is organized to support life.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →