← Latest papers
🧬 biology

GENEB: Why Genomic Models Are Hard to Compare

To address the difficulty of comparing genomic foundation models due to fragmented benchmarks and incompatible protocols, the paper introduces GENEB, a unified large-scale diagnostic benchmark that evaluates 40 models across 100 tasks to reveal the instability of aggregate rankings and demonstrate that architectural and pretraining alignment often outweigh model scale.

Original authors: Daria Ledneva, Mikhail Nuridinov, Denis Kuznetsov

Published 2026-08-06
📖 5 min read🧠 Deep dive

Original authors: Daria Ledneva, Mikhail Nuridinov, Denis Kuznetsov

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are trying to teach a computer to read the secret instruction manual of life. This manual is written in a language made of just four letters: A, C, G, and T. These letters form the DNA code that tells every living thing—from a tiny bacterium to a giant oak tree—how to build itself and function. In recent years, scientists have built "foundation models," which are like super-smart students that have read billions of pages of this DNA manual. They are supposed to learn the rules of the language so well that they can predict things like how a gene will behave or if a specific sequence causes a disease.

But here is the problem: imagine if every student in a class took a different test, used a different dictionary, and was graded by a different teacher. One student might be great at spelling but terrible at grammar, while another is a grammar wizard but can't spell. If you just looked at their final grades, you wouldn't know who is actually the best student overall. You wouldn't know if the "grammar wizard" failed the spelling test because they are bad at it, or because the test was unfair. This is exactly what has been happening in the world of DNA AI. Different research teams build their models, test them on their own unique sets of tasks, and claim their model is the "best." But because they aren't taking the same test, nobody can really compare them fairly. It's like trying to decide who is the best athlete in the world by comparing a swimmer's time in a pool to a runner's time on a track without a unified standard.

This is where a new study called GENEB steps in to fix the mess. The researchers, Daria Ledneva, Mikhail Nuridinov, and Denis Kuznetsov, decided to build a massive, unified "Olympics" for DNA models. They gathered 40 different genomic foundation models—some huge, some small, some trained on human DNA, others on bacteria or plants—and put them all through the exact same 100 different challenges. These challenges covered 13 different types of biological tasks, like finding specific switches in the DNA (promoters), identifying how genes are turned on or off (histone modifications), or spotting viruses. They used a strict, fair scoring system to see how each model performed when frozen in place, meaning they couldn't learn anything new during the test; they had to rely entirely on what they had already learned.

The results of this grand experiment were surprising and shook up a few common beliefs. First, the study found that bigger isn't always better. In the world of AI, there's a popular idea that if you just make the model bigger (give it more "brain power" or parameters), it will automatically get smarter. But GENEB showed that this isn't a magic bullet. They found that a tiny model with only 86 million "parameters" (a measure of size) could actually beat a giant model with 7 billion parameters by a huge margin. In fact, the giant model was so bad at some tasks that it performed worse than random guessing, simply because it had been trained on the wrong kind of DNA (bacteria) while the test was about human DNA. It's like taking a chef who is a master at cooking Italian food and asking them to cook a traditional Japanese meal; no matter how famous or expensive the chef is, they will struggle if they don't know the ingredients.

Second, the study discovered that what a model was trained on matters more than how big it is. A model trained specifically on human DNA or a mix of many different species performed much better on human tasks than a massive model trained only on bacteria. The researchers showed that if the "diet" of the model (the data it ate during training) doesn't match the "meal" it's being asked to eat later (the task it's tested on), it fails. For example, models trained on human data were terrible at identifying plant genes, and models trained on bacteria couldn't figure out how human genes are spliced together. The paper suggests that for these models to be truly useful, they need to be fed the right kind of data for the job, rather than just being made larger.

Finally, the paper argues that there is no single "champion" model. Just like in sports, a model might be the best at running but terrible at swimming. The study found that a model that was the top performer on one task could be near the bottom on another. This means that instead of looking for one "best" model to rule them all, scientists and doctors need to pick the right tool for the specific job. If you need to predict how a human gene is regulated, you should pick a model trained on human epigenetic data. If you need to identify plant genes, you need a model trained on plants. The paper concludes that the current way of ranking these models with a single leaderboard is misleading and that we need a more careful, category-by-category approach to choose the right AI for the right biological question.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →