← Latest papers
🧬 biology

Leveraging Structural Variants to Unlock Hidden Genetic Diversity for Tomato Crop Improvement

This study demonstrates that repurposing cost-effective short-read sequencing data from diverse tomato lines can uncover over 71,000 high-confidence structural variants, revealing them as critical, previously hidden drivers of genetic diversity that enhance breeding workflows and resolve lineage relationships without requiring additional long-read sequencing investments.

Original authors: Reza Shekasteband

Published 2026-08-21
📖 5 min read🧠 Deep dive

Original authors: Reza Shekasteband

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Tomatoes are one of the world's most important vegetables, and for decades, scientists have studied their DNA to understand why some plants resist disease, why others produce larger fruits, and how different varieties are related. For a long time, researchers focused on the smallest changes in the genetic code, looking at single letters that differ between plants. These tiny differences, known as single-nucleotide variants, have been incredibly useful for breeding better crops. However, the genome also contains much larger changes, such as missing chunks of DNA, extra copies of sections, or pieces that have moved to new locations. These larger alterations are called structural variants. While they are known to be powerful drivers of how a plant looks and behaves, finding them has traditionally been difficult and expensive, often requiring specialized, high-cost technology that many breeding programs cannot afford.

A new study by Reza Shekasteband at North Carolina State University shows that scientists can now find these hidden, large-scale genetic changes using data that already exists. The researcher took advantage of a massive collection of tomato DNA sequences that had been generated over six years using standard, affordable sequencing machines. Instead of running new, expensive tests, the team applied a specific computer method to this existing data to hunt for the larger structural changes. They examined sixty different tomato lines, ranging from wild ancestors growing in the Andes to modern, commercial varieties grown in fields. By re-analyzing the short DNA reads already in the database, the team successfully identified over 71,000 high-confidence structural variants. This approach proves that valuable genetic information is sitting idle in old datasets, waiting to be unlocked without spending a single dollar on new sequencing.

The study revealed that these large genetic changes are not spread evenly across the tomato genome. Instead, they cluster in specific areas, particularly near the ends of the chromosomes and in regions known to control disease resistance. When the researchers compared the wild tomatoes to the cultivated ones, they found that the wild plants carried nearly twice as many of these structural changes. In the wild genomes, missing pieces of DNA were the most common type of change, whereas in the cultivated tomatoes, extra insertions of DNA were more frequent. This difference highlights how domestication has shaped the genetic architecture of the crop, filtering out some of the wild variation while introducing new structural features. The team also discovered that many modern breeding lines carry unique structural changes that appear nowhere else, acting as a genetic fingerprint for specific varieties.

One of the most significant findings was that these large structural changes tell the same evolutionary story as the tiny single-letter changes, but with greater clarity. When the researchers built family trees based on the structural variants, the results matched the trees built from the tiny changes almost perfectly. However, the structural data was better at sorting out confusing relationships. For example, one wild tomato variety that had been grouped with cultivated grapes in the old analysis was correctly placed back with its wild relatives when the structural changes were considered. This suggests that looking at the larger pieces of DNA helps breeders understand the true history of their plants and make smarter decisions about which varieties to cross.

The research also uncovered a practical lesson for how scientists process this data. The team found that the way they aligned the DNA reads to the reference genome made a huge difference. A standard method that forces reads to match perfectly from end to end failed to find any structural changes. By switching to a more flexible approach that allows reads to match partially, the team could detect the breaks and overlaps that signal a structural variant. This adjustment allowed them to find thousands more changes than they would have otherwise. They tested different ways of grouping the samples and found that analyzing batches of diverse lines together yielded the most comprehensive results, revealing unique changes in specific lines that were missed when looking at samples individually.

Despite the success, the study acknowledges that this method has limits. While it is excellent at finding deletions and insertions, it may miss very complex rearrangements or the precise structure of inserted genes, such as those used in genetically modified plants. The researchers confirmed this by finding that their method missed some insertion sites in transgenic lines that other techniques could see. Nevertheless, the study demonstrates that a cost-effective, scalable strategy exists for mining archived data. By using standard short-read sequencing data and optimizing the computer tools, breeding programs can now access a hidden layer of genetic diversity. This allows them to track specific traits, refine their understanding of plant families, and accelerate the development of new tomato varieties without the prohibitive cost of new, high-end sequencing technologies. The work turns what was once considered a limitation of older data into a powerful resource for the future of crop improvement.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →