← Latest papers
🧬 biology

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

JEPA-DNA introduces a model-agnostic continual training framework that integrates Joint-Embedding Predictive Architectures with traditional generative objectives to shift genomic foundation models from token-level reconstruction to semantic alignment, achieving state-of-the-art performance across diverse genomic benchmarks.

Original authors: Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Nicole Bussola, Simon Lee, Shane O'Connell, Dung Hoang, Mariss
Published 2026-08-12
📖 3 min read☕ Coffee break read

Original authors: Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Nicole Bussola, Simon Lee, Shane O'Connell, Dung Hoang, Marissa Wirth, Alexander W. Charney, Nati Daniel, Yoli Shavit

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine the human body as a massive, ancient library. Inside this library, every single instruction for building and running a person is written in a code made of just four letters: A, C, G, and T. This code is DNA. For a long time, scientists have been trying to teach computers to read this library, hoping that if a computer understands the "grammar" of DNA, it can predict how our bodies work, why we get sick, or how to fix genetic glitches.

To do this, researchers use something called "Foundation Models." Think of these as super-smart students who have read millions of books (DNA sequences) and learned the rules of the language. Usually, these students learn by playing a game of "fill in the blank." If you cover up a word in a sentence, the student has to guess what it was based on the words around it. This works great for learning the local rules of the sentence, like how words fit together. But here's the problem: knowing the grammar doesn't always mean you understand the story. A student might know that "The cat sat on the..." is usually followed by "mat," but they might not understand that the cat is actually hiding from a dog, or that the whole scene is part of a larger mystery. In DNA, this means computers are good at spotting small patterns but often miss the big picture of how different parts of the code work together to control life.

This is where a new study called JEPA-DNA comes in. The researchers, a team from NVIDIA and medical centers in Israel and the US, decided to upgrade these DNA-reading students. They introduced a new training method that forces the computer to stop just guessing the missing letters and start guessing the meaning of the missing parts. Instead of asking, "What letter goes here?" they ask, "What is the job of this whole chunk of DNA?"

The team tested this new method on 17 different tasks, ranging from finding where genes start (promoters) to predicting how genetic changes affect health. They tried it on three different types of DNA-reading computers (DNABERT-2, NTv3, and HyenaDNA) to make sure it worked for everyone. The results were impressive: by adding this new "meaning-focused" training, the computers got better at almost every task. For example, on one task called "GUE Promoter," the new method improved performance by over 11%. On another task involving disease variants, it boosted the ability to spot problems by more than 12%.

The paper suggests that this approach works because it teaches the computer to look at the "forest" instead of just the "trees." While the old methods were great at memorizing the local syntax (the specific order of letters), JEPA-DNA helps the model understand the global functional context (what that sequence actually does in the body). The researchers found that this new way of learning creates a "world model" of biology, where the computer understands the rules of life, not just the spelling. They even showed that this method works better than the current best models available, setting a new high bar for what these AI systems can achieve. It's like taking a student who was good at spelling and teaching them to be a literary critic who understands the plot, the characters, and the theme all at once.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →