← Latest papers
🧬 biology

A Computational Cancer GWAS Screen Highlights Structural Variation at the 7q22.1 Locus Near STAG3

This study demonstrates that a computational screen integrating short-read and long-read sequencing data reveals a significant enrichment of structural variants at meiosis-specific gene loci overlapping cancer GWAS signals, specifically identifying a novel 442 bp deletion near the *STAG3* gene as a candidate driver at the 7q22.1 colorectal cancer susceptibility locus that would have been missed by short-read sequencing alone.

Original authors: Koushik Chowdhury

Published 2026-09-18
📖 7 min read🧠 Deep dive

Original authors: Koushik Chowdhury

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

To understand how cancer begins, scientists often look at the instructions written in our DNA. These instructions are not just a list of single letters; they are a complex library where entire paragraphs can be missing, duplicated, or rearranged. For decades, researchers have scanned these libraries looking for tiny typos—single letter changes—to find clues about why some people develop cancer while others do not. However, this method has a blind spot. It often misses the larger, more dramatic rearrangements of the genetic text, known as structural variants. These are chunks of DNA, ranging from fifty letters to thousands, that are deleted, copied, or moved. Because these changes often happen in messy, repetitive regions of the genome, the standard tools used to read DNA have struggled to see them clearly. This gap is particularly wide when looking at genes that are active during meiosis, the special cell division process that creates sperm and eggs. These genes are vital for repairing DNA, and when they malfunction, they can leave the door open for cancer to develop later in life.

A new study by Koushik Chowdhury at Saarland University in Germany has taken a fresh look at these hidden regions. By combining two powerful sets of genetic data—one built from standard short-read sequencing and another from advanced long-read technology that can see through the messy parts of the genome—the researcher created a detailed map of structural variants near thirty-one genes involved in meiosis. The goal was to see if these large genetic rearrangements, which had been largely ignored in previous cancer studies, actually line up with known cancer risk spots. The study found that these structural variants are not random noise; they are significantly more likely to be found near cancer risk signals than would be expected by chance. Most notably, the research pinpointed a specific, 442-letter deletion on chromosome 7 that sits right next to a known risk factor for colorectal cancer. This deletion is located near a gene called STAG3, which acts as a backup version of a gene already known to drive colorectal cancer when it malfunctions in the body. While the study does not prove that this deletion causes cancer, it suggests that STAG3 is a strong candidate for further investigation, offering a new lead on how inherited genetic changes might influence cancer risk.

The study began by acknowledging a limitation in how cancer genetics has been practiced for years. Standard genetic tests, which use arrays to scan the genome, are excellent at finding single-letter changes but are terrible at spotting larger structural rearrangements. Imagine trying to read a book where entire pages have been torn out or pasted in the wrong order; a standard scanner might only see the text on the remaining pages and miss the missing or misplaced sections entirely. This is especially true for genes involved in meiosis, which are often surrounded by repetitive DNA sequences that confuse standard reading tools. To fix this, Chowdhury used two different datasets. The first came from gnomAD, a massive database of genetic information from about 63,000 people, which was generated using standard short-read sequencing. The second came from the Human Pangenome Reference Consortium, which used long-read sequencing technology to map the genomes of 47 individuals with much greater precision, capturing complex regions that the first dataset missed.

The researcher focused on thirty-one specific genes. Seventeen of these were well-known cancer genes, while the other fourteen were meiosis-enriched genes, meaning their primary job is to function during the creation of sperm and eggs. The study asked a simple question: do structural variants in these genes overlap with known cancer risk locations? To answer this, the researcher compared the genetic maps against a catalog of over 6,000 cancer risk signals identified in previous studies. The results were striking. Structural variants found in the windows around these thirty-one genes were nearly eight times more likely to overlap with a cancer risk signal than variants found elsewhere in the genome. This suggests that the blind spot in previous research was real and that these larger genetic changes are likely playing a role in cancer susceptibility that scientists had been missing.

One of the most significant findings was that the long-read data from the pangenome project revealed 187 structural variants that the standard short-read data completely failed to detect. These "invisible" variants were not scattered randomly; they were concentrated in the meiosis-enriched gene regions. This confirms that the complex, repetitive nature of these specific genes makes them difficult to study with older technology, and that the new long-read tools are essential for seeing the full picture. The study identified 913 total associations involving these meiosis-enriched genes, linking them to various cancers including breast, prostate, and skin cancer.

The most detailed part of the study focused on a specific location on chromosome 7, known as the 7q22.1 locus, which is a known hotspot for colorectal cancer risk. Previous studies had identified two single-letter changes in this area that were strongly linked to the disease, but because these changes sat in a region between genes, scientists did not know which gene they were affecting. The standard maps pointed to a nearby gene called TRIM4 simply because it was the closest neighbor, but there was no evidence that TRIM4 was actually involved in the disease.

Chowdhury's analysis uncovered a 442-letter deletion sitting right between the two known risk signals. This deletion was found in about 0.9 percent of the population. While the study could not definitively prove that this deletion is the direct cause of the risk, its position is highly suspicious. It sits in the same small genetic neighborhood as the risk signals, within a window of just 8,700 letters. The researcher then looked at the genes in this area to see which one might be the true target. Two genes, GJC3 and AZGP1, sit physically between the deletion and the next major gene, but neither showed any signs of being regulated by the risk signals in colon tissue. However, the gene STAG3, located about 289,000 letters downstream, showed a strong signal. STAG3 is the meiosis-specific version of STAG2, a gene that is frequently mutated in colorectal cancer tumors. The study found that STAG3 has a known link to gene expression in the colon, suggesting that the risk signals and the deletion might be regulating STAG3 from a distance.

The paper is careful not to claim that this discovery solves the mystery of colorectal cancer. The researcher notes that the deletion and the risk signals have not been formally proven to be linked in the same DNA strand, and no lab experiments have been done to test if the deletion actually changes how STAG3 works. The study is a computational screen, a way of sorting through vast amounts of data to find the most promising leads. It suggests that STAG3 is a candidate gene that deserves to be studied further, particularly because it fits a biological story: a gene that helps repair DNA during meiosis, when it is inherited in a damaged form, could make a person more susceptible to cancer later in life.

This work highlights a shift in how scientists approach cancer genetics. By moving beyond simple letter-by-letter scans and embracing technologies that can see the larger structural changes, researchers are finding new connections between our inherited DNA and disease. The study demonstrates that the "missing" structural variants are not just noise; they are likely key pieces of the puzzle, especially in the complex regions of the genome that standard tools have struggled to read. For the 7q22.1 locus, the findings provide a concrete direction for future research, moving the focus from a generic nearby gene to a specific, biologically plausible candidate that could explain why some people are born with a higher risk of developing colorectal cancer.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →