Position-resolved k-mer analysis of mature microRNAs reveal widespread CG and UA depletion across taxa
This study analyzes nearly 49,000 mature microRNA sequences across 271 organisms to reveal a widespread, lineage-associated positional organization characterized by significant depletion of CG and UA dinucleotides and enrichment of UG, suggesting structural or evolutionary constraints beyond simple mutational bias.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Inside every living cell, a vast network of tiny molecular instructions keeps the machinery of life running smoothly. Among the most important of these instructions are short strands of RNA called microRNAs. These molecules act as precise regulators, finding specific matching sequences in other genetic messages and silencing them to ensure the cell produces the right proteins at the right time. For a microRNA to work, its shape and chemical makeup must be exact. Scientists have long known that the very first few letters of a microRNA's sequence are critical, acting like a key that fits into a specific lock to start the process. However, the rules governing the rest of the molecule—the letters that follow those initial few—have remained a mystery. Do the letters that sit next to each other in these tiny strands follow a hidden pattern, or are they arranged randomly?
A team of researchers from Bharathidasan University in India has now mapped the entire landscape of these tiny molecules to answer that question. They did not look at just a handful of examples; instead, they examined nearly 49,000 records of mature microRNAs from 271 different organisms, ranging from humans and other mammals to insects, fish, and plants. By analyzing this massive collection, they discovered that the arrangement of letters in these molecules is far from random. There is a widespread, consistent rule that governs how these short genetic strands are built, a rule that holds true across the entire tree of life.
The researchers focused on how often specific pairs of letters appear next to each other. In the language of genetics, these pairs are called dinucleotides. When they counted every possible pairing across the millions of sequences they studied, two specific pairs stood out for their absence. The combination of the letters C and G, and the combination of U and A, appeared far less often than chance would predict. In fact, the C-G pair was the rarest of all sixteen possible combinations, making up only about 3.3 percent of the total. The U-A pair was the second rarest. Conversely, the pair U-G was the most common, appearing significantly more often than expected.
To ensure these findings were not just a result of the numbers, the scientists used a rigorous method of comparison. They created two different sets of "what-if" scenarios. In the first, they took the exact same microRNA sequences and scrambled the order of the letters within each one, keeping the total number of each letter the same but destroying any natural order. In the second, they shuffled the letters at each specific position across all the sequences of the same length. Even when compared against these scrambled, random versions, the real microRNAs still showed a distinct lack of C-G and U-A pairs. This proved that the scarcity of these pairs is a genuine feature of the molecules, not just a side effect of how many of each letter happens to be present.
This pattern of scarcity was not limited to just the two-letter pairs. The researchers also looked at groups of three letters, known as trinucleotides. They found that every single combination containing a C-G pair or a U-A pair was also depleted, appearing less frequently than the random models predicted. This suggests that the rule against these specific neighbors applies throughout the entire length of the microRNA, not just in one specific spot. The finding was consistent across all ten major groups of organisms they tested, from flowering plants to nematode worms and humans. Whether the organism was a plant or an animal, the microRNAs consistently avoided placing a C next to a G or a U next to an A.
There was one interesting exception that required closer inspection. While the U-A pair was rare overall, the researchers noticed a small spike in its frequency at the very beginning of the microRNA sequence. At the first position, the U-A pair appeared more often than in the rest of the molecule. However, when the scientists accounted for the fact that the first letter of a microRNA is almost always a U, this spike disappeared. It turned out that the high number of U-A pairs at the start was simply because there were so many U's at the start to begin with, not because the molecule had a special preference for placing an A right after a U. Once this initial composition was taken into consideration, the U-A pair was found to be depleted everywhere, just like the C-G pair.
The study also explored why these patterns might exist. For the missing C-G pairs, a plausible explanation lies in the history of DNA mutation. In many animals, the letter C is often chemically modified in a way that makes it prone to changing into a T (or U in RNA) over evolutionary time. If a C-G pair in the DNA mutates into a U-G pair, the resulting RNA would naturally have fewer C-G pairs and more U-G pairs. This matches the researchers' observation that U-G is the most common pair while C-G is the rarest. However, the paper notes that this is only a hypothesis; the study did not directly measure DNA methylation or mutation rates, so this remains a likely possibility rather than a confirmed cause.
The reason for the scarcity of the U-A pair, however, remains a complete mystery. The researchers could not find a biological mechanism that explains why these two letters avoid each other. It is possible that the structure of the molecule itself, the way it is processed inside the cell, or the way it interacts with its targets makes the U-A combination unstable or undesirable. The study successfully identified that this depletion exists and is universal, but it stops short of explaining why.
By creating a detailed map of how letters are arranged in these tiny regulators, the researchers have provided a new foundation for understanding microRNA biology. They have shown that these molecules are not just random strings of letters but are built with a specific, conserved architecture that has been maintained across millions of years of evolution. While the full reason for this architecture is not yet known, the discovery that microRNAs universally avoid certain letter combinations offers a new clue for scientists trying to understand how these critical molecules function and evolve. The work suggests that the rules of life are written not just in the individual letters, but in the specific ways those letters are allowed to sit next to one another.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.