← Latest papers
💻 bioinformatics

Scanning transcriptomes for nonlinear, domain-level similarities using hmSEEKR

The paper introduces hmSEEKR, a k-mer-based hidden Markov model that scans transcriptomes to identify non-linear, domain-level sequence similarities in long noncoding RNAs, thereby enabling the discovery of functionally related RNA domains and their associated protein interaction networks without requiring prior knowledge of sequence alignment.

Original authors: Li, S., Sprague, D. A., Eberhard, Q. E., Boyson, S. P., Laederach, A., Calabrese, J. M.

Published 2026-07-08
📖 5 min read🧠 Deep dive

Original authors: Li, S., Sprague, D. A., Eberhard, Q. E., Boyson, S. P., Laederach, A., Calabrese, J. M.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

The Big Problem: Finding a Needle in a Haystack (That Doesn't Look Like a Needle)

Imagine you are trying to find a specific recipe for a cake in a giant library of cookbooks. Usually, you would look for the recipe by searching for the title or the list of ingredients (the "linear sequence").

But long non-coding RNAs (lncRNAs) are like abstract art recipes. They don't have a standard title, and their ingredients might be listed in a completely different order than other recipes that make the exact same cake. Because they look so different on paper, traditional search tools can't find them. Scientists know these RNAs are important for running the cell's "factory," but they can't figure out what most of them actually do because they can't find similar ones to compare them to.

The Solution: hmSEEKR (The "Smell" Detector)

The authors built a new tool called hmSEEKR. Instead of looking for the exact order of ingredients, hmSEEKR looks for the overall "flavor profile" or the "aroma" of the recipe.

  • The Analogy: Imagine you are trying to find a specific type of coffee. A traditional search looks for the exact brand name on the bag. hmSEEKR, however, sniffs the air. It knows that a "dark roast with hazelnut notes" has a specific scent profile. Even if the coffee is in a different bag, written in a different language, or mixed with other beans, hmSEEKR can sniff it out and say, "Hey, this smells just like that hazelnut coffee!"

How It Works: The "Two-State" Detective

The tool uses a statistical method called a Hidden Markov Model (HMM). Think of this as a detective walking through a long hallway of text, deciding at every single step: "Is this part of the 'Target' room, or is it just the 'Background' wall?"

  1. The Query (The Target): You give the tool a specific piece of RNA you know works (like a specific domain in the famous XIST RNA). This is your "Target Room."
  2. The Null (The Background): The tool also knows what "normal" RNA looks like (the "Background Wall").
  3. The Scan: The tool scans the entire library of RNA. It breaks the text into tiny chunks (called k-mers).
    • If a chunk smells like the "Target Room," the detective marks it as a Hit.
    • If it smells like the "Background Wall," it ignores it.
  4. The Result: It connects the dots. If it finds a long stretch of "Target Room" smells, it flags that section as a match, even if the rest of the RNA looks totally different.

What They Discovered

The team tested this tool using three famous RNAs: XIST, NEAT1, and MALAT1. These are like the "Superstars" of the RNA world because we know exactly what they do.

1. It Can Find Hidden Copies
They took pieces of these Superstar RNAs, chopped them up, and hid them randomly inside other RNA sequences. hmSEEKR found 100% of them, even when the hidden pieces were mutated (changed slightly). It proved the tool is very good at spotting the "flavor" even when the "ingredients" are slightly off.

2. It Finds Functional Twins
When the tool found a match in a different RNA, they checked if that match acted like the original. They looked at which proteins (the cell's workers) stuck to these new matches.

  • The Result: The proteins that stuck to the new matches were the same proteins that stuck to the original Superstar RNAs. This suggests the new matches are likely doing the same job as the originals.

3. The "Minimalist" Search
They asked: "Can we find RNAs that look like a whole XIST or NEAT1, just with the essential parts in the right order?"

  • They found hundreds of other RNAs that had the "XIST recipe" or the "NEAT1 recipe" scattered through them in the correct sequence.
  • The Twist: Most of these "look-alikes" were actually nascent transcripts (RNA that is still being written and hasn't been finished yet) or came from protein-coding genes, not just the usual lncRNA suspects. This suggests that the "factory floor" (where RNA is made) is full of these regulatory tools.

4. The "Activator" vs. "Repressor" Test
Finally, they looked at RNAs known to either turn genes ON (activators) or turn them OFF (repressors).

  • They found that RNAs meant to turn things OFF tended to have "flavors" similar to known "repressor" domains (like those that bind to HNRNP proteins).
  • RNAs meant to turn things ON had "flavors" similar to "activator" domains (like those that bind to Mediator or P300 complexes).
  • The Takeaway: You can guess what an RNA does just by sniffing its "flavor profile."

Summary

hmSEEKR is a new search engine for the cell's instruction manual. It doesn't care if the sentences are written in the same order; it cares if the vibe is the same. By doing this, it helps scientists find hidden functional parts of RNA that were previously invisible, showing us that the cell is full of "functional twins" that look different on paper but do the exact same job.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →