← Latest papers
🧬 biology

Copy aware deconvolution of Hi-C interactions in extrachromosomal DNA using CADET

The paper introduces CADET, a computational method that directly infers copy-specific Hi-C interaction matrices for extrachromosomal DNA (ecDNA) using expectation-maximization, thereby resolving the ambiguity caused by duplicated genomic segments without requiring three-dimensional structural reconstruction.

Original authors: Biswanath Chowdhury

Published 2026-09-24
📖 6 min read🧠 Deep dive

Original authors: Biswanath Chowdhury

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Inside the nucleus of a cancer cell, the genetic blueprint is often in a state of chaos. While healthy cells keep their DNA neatly organized into long, linear chromosomes, cancer cells frequently break these strands apart and reassemble them into circular loops known as extrachromosomal DNA. These loops are not just random tangles; they act as powerful engines for cancer growth, often carrying extra copies of genes that drive the disease forward. Because these circular molecules can hold multiple copies of the same genetic segment, they create a unique problem for scientists trying to map the cell's internal architecture. When researchers use high-throughput techniques to see how different parts of the DNA touch and interact, the signals from these duplicate copies merge into a single, blurry image. It is as if looking at a photograph where three identical twins have been superimposed; you can see the group, but you cannot tell which twin is doing what.

To understand how these cancer cells function, scientists need to separate these overlapping signals. They need to know which specific copy of a gene is interacting with which specific part of the genome to drive the disease. A new method called CADET, developed by researcher Biswanath Chowdhury, offers a way to do exactly this. Instead of trying to build a three-dimensional model of the DNA shape to guess where the copies are, this approach works directly with the contact data. It treats the problem like untangling a knot of overlapping voices. By using the known layout of the circular DNA and looking for subtle differences in how the DNA copies interact with their immediate neighbors, the method can mathematically separate the merged signals. This allows researchers to see the distinct interactions of each individual copy, revealing a clearer picture of the regulatory architecture that fuels the cancer.

The researchers tested this new approach using computer simulations first, creating fake datasets where they knew the exact truth about how the DNA copies were interacting. They introduced varying levels of duplication, sometimes making the copies behave identically and other times giving them different behaviors. In every scenario, the method successfully recovered the hidden, copy-specific interactions from the blurred data. It proved particularly effective when the duplicated copies had different local environments, using those subtle differences to tell them apart. The simulations showed that the method could accurately identify which copy was responsible for a specific contact about 80 to 83 percent of the time in complex situations, and it did so without changing the total number of contacts observed in the original data. This confirmed that the method was not inventing new information but rather redistributing existing information into a more useful format.

The team then applied CADET to real-world data from three different cancer cell lines, each with its own unique arrangement of extrachromosomal DNA. In one case, a lung cancer cell line, the DNA was highly duplicated, with large sections of the genome repeated multiple times. The method successfully untangled these repetitions, revealing that certain parts of the DNA were acting as hubs, repeatedly interacting with different partners. In another case, a brain cancer cell line, the circular DNA brought together segments from two different chromosomes. Here, the method identified a specific, strong interaction between a cancer-driving gene and a distant regulatory element on a different chromosome, a finding that matched previous studies but was now resolved without needing to reconstruct the 3D shape of the molecule. In a third case involving a neuroblastoma cell line, the DNA was arranged in a tandem repeat on a chromosome rather than a circle. Even though the individual signals were weaker in this dataset, the method still found recurring patterns of interaction around specific genes known to be involved in the disease.

What makes this approach distinct is how it handles the ambiguity of duplicated DNA. Previous methods often tried to solve this by building a 3D model of the DNA structure first, assuming that the shape of the molecule dictated how the signals should be split. CADET, however, skips that step. It looks directly at the contact map and uses statistical reasoning to decide how to split the signal based on the distance between segments and the local support each copy receives. This means the resulting map of interactions is independent of any specific 3D shape assumption. The researchers found that while the method did not always pinpoint the exact strength of every single interaction, it was very good at identifying families of interactions that were enriched or stronger than expected. In the brain cancer cell line, for instance, the method identified a module of interactions involving a key cancer gene that matched findings from other studies, but it placed this interaction within a context of recurring contact patterns rather than a single structural event.

The study also highlighted the limitations of what can be known from this type of data. When two copies of a DNA segment are in identical local environments, the method cannot distinguish between them, just as a listener cannot tell two identical voices apart if they are speaking the same words at the same volume. In these cases, the method assigns the signal equally, acknowledging that the data does not contain enough information to make a finer distinction. The researchers were careful to note that while they could identify candidate interactions and group them into functional modules, these findings represent hypotheses that need to be tested with further experiments. They did not claim to have solved the problem of cancer gene regulation, but rather provided a new, clearer lens through which to view the complex, tangled interactions within cancer cells.

By separating the merged signals of duplicated DNA, this work offers a more direct way to study the regulatory networks that drive cancer. It allows scientists to see which specific copies of a gene are active and how they are communicating with the rest of the genome. The method has already revealed distinct patterns of interaction in different types of cancer, showing that these circular DNA molecules organize themselves in diverse ways. Whether the DNA is arranged in a circle or a long repeat, the ability to resolve these copy-specific contacts opens the door to a deeper understanding of how cancer cells rewire their genetic instructions to survive and grow. This clarity could eventually help researchers identify new targets for therapy, focusing on the specific interactions that are most critical to the disease.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →