← Latest papers
📄 plant biology

BOTANIC-1: a series of long-context plant genomic foundation models in the agentic era

This paper introduces Botanic1, a family of long-context, agent-integrated genomic language models trained on unannotated plant data that outperform existing models in predicting trait-associated regions and provide biological insights through mechanistic interpretability, all while being made available to the research community to accelerate climate-resilient crop development.

Original authors: Barozet, A., Cabeli, V., Ogier du Terrail, J., Rukhovich, A., Janssoone, T., Klajer, G., Sheikhitarghi, Z., Andrews, G., Veran, C., Strouk, L.

Published 2026-09-07
📖 5 min read🧠 Deep dive

Original authors: Barozet, A., Cabeli, V., Ogier du Terrail, J., Rukhovich, A., Janssoone, T., Klajer, G., Sheikhitarghi, Z., Andrews, G., Veran, C., Strouk, L.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine a world where the code of life is written in a language of four letters: A, C, G, and T. These letters form the DNA sequences that dictate how a plant grows, how it resists drought, and how it fights off disease. For decades, scientists have tried to read this language, but it is a massive, complex text with billions of words, much of it written in a dialect that changes from one species to another. When a plant faces a harsh climate or a new pest, the solution often lies in a tiny change to this code—a single letter swap that might make a crop stronger. Finding that specific change is like searching for a needle in a haystack, but the haystack is the size of a library, and the needle is invisible to the naked eye. Traditional methods rely on comparing many plants to find patterns, but this is slow, expensive, and often leaves researchers with thousands of possibilities instead of one clear answer.

A team of researchers in Paris has now introduced a new way to read this biological text. They have built a family of artificial intelligence models, which they call Botanic1, designed specifically to understand the grammar of plant DNA. Unlike previous tools that needed to be taught the rules of biology by humans, these models learned the language on their own. They were fed the genetic sequences of 320 different plant species, ranging from tiny mosses to towering trees, and asked to predict missing letters in the sequence. By doing this millions of times, the models learned the hidden rules of how DNA is structured, how genes are turned on and off, and how tiny changes in the code affect the plant. The result is a system that can look at a long stretch of DNA and instantly sense which parts are critical for the plant's survival and which changes might be harmful or beneficial.

The researchers did not just build one model; they created a "factory" that uses intelligent software agents to manage the entire process of training and testing these models. These agents handle the heavy lifting, organizing data, running experiments, and fixing errors, allowing human scientists to focus on the big ideas. This approach allowed them to train models that can read DNA sequences up to 128,000 letters long, a length that captures the long-range interactions between different parts of the genome that were previously impossible to analyze together. In tests, these models outperformed every other existing tool for understanding plant genetics, including those that were much larger and trained on far more data. They were able to identify the most important parts of the DNA, such as the boundaries where genes start and stop, and even spot the specific signals that tell a cell how to cut and paste its genetic instructions.

What makes this discovery particularly powerful is that the models learned these rules without being told what they were. When the researchers looked inside the models to see what they had learned, they found that the AI had independently discovered biological concepts like "splice sites," which are the instructions for how a gene is edited before it becomes a protein. In one striking example, the model identified a specific location in a gene that the standard scientific databases had marked incorrectly for over twenty years. The model's internal logic pointed to a different spot, and when the researchers checked the history of that gene, they found that the original 1999 description was actually correct, and the model had rediscovered the truth that had been lost in later updates. This suggests that these models can act as a new kind of microscope, revealing biological truths that are hidden in the data but missed by human annotation.

The team also demonstrated how these models can work alongside human researchers in a new way. They set up a scenario where a general-purpose artificial intelligence agent was asked to find the cause of a specific trait in melons: why some plants produce female flowers and others produce both male and female flowers. The agent was given a list of thousands of genetic differences and asked to find the one responsible. When the agent tried to solve this using only standard tools, it could narrow the list down but could not pick the right answer. However, when the researchers gave the agent access to the Botanic1 model as a tool, the agent immediately ranked the correct genetic change as the number one candidate. The model provided a continuous measure of how surprising a genetic change was, which helped the agent distinguish the true cause from the noise.

This work represents a significant step forward in how we study plants. By teaching machines to read the language of DNA, the researchers have created tools that can accelerate the search for crops that can withstand a changing climate. These models are not just faster; they are more accurate and can see patterns that were previously invisible. They offer a way to move from guessing which genes might be important to knowing with high confidence which ones are. As the researchers release these tools to the scientific community, they open the door for a new era of plant biology, where the complex code of life can be understood and improved with a speed and precision that was previously out of reach. The models are now available for other scientists to use, promising to help breed better crops and understand the fundamental mechanisms of life in the plant kingdom.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →