Flow Orchestrated Regulatory Genomics Engine (FORGE): A Configurable Nextflow Pipeline for End-to-End snMultiome Analysis
FORGE is a configurable, containerized Nextflow pipeline that automates end-to-end single-nucleus multiome analysis by integrating gene expression and chromatin accessibility data for regulatory network inference and differential testing, thereby enhancing scalability, reproducibility, and biological interpretation across diverse datasets.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Inside every cell of our bodies, a complex conversation is constantly taking place. One part of the cell, the nucleus, holds the master instruction manual, the DNA, which contains the genes that determine what a cell is and what it does. Another part of the nucleus, the chromatin, acts like a set of switches and dimmers that decide which of those instructions are turned on or off at any given moment. In a healthy body, these switches work in perfect harmony, allowing a skin cell to stay a skin cell and a brain cell to function as a brain cell. However, when this regulatory system goes wrong, it can lead to diseases like Alzheimer's, where the brain's cells lose their ability to communicate and function properly. For a long time, scientists could only look at the instruction manual or the switches separately, making it difficult to see how they worked together to control the cell's behavior.
A new tool called FORGE, developed by researchers at the University of California, Irvine, changes how we can study this relationship. This tool is a sophisticated computer program designed to analyze a specific type of biological data called "multiome" data. This data is special because it captures both the gene activity and the chromatin switches from the exact same tiny cell nucleus at the same time. Before this tool existed, scientists had to stitch together different, often incompatible computer programs to make sense of this data, a process that was slow, prone to errors, and difficult to repeat. FORGE automates this entire process, taking raw data from a cell and guiding it through a series of checks and analyses to reveal the hidden rules that control how genes are turned on and off. It acts like a universal translator, ensuring that the story told by the genes matches the story told by the switches, and doing so in a way that is reliable enough for other scientists to trust and use.
The researchers tested this tool on four different sets of biological data, ranging from human blood cells to mouse brain and kidney tissue. They wanted to see if FORGE could handle different types of cells, different species, and different experimental setups. In one major test, they applied it to a dataset of human blood cells. The tool successfully identified the different types of immune cells present, such as T cells and monocytes, and mapped out the specific genetic programs that drive their behavior. It found that the tool could agree with other established methods on what these cells were, but it went further by uncovering new details about how specific proteins, called transcription factors, bind to the DNA to control cell function. For instance, it highlighted a specific group of proteins that seemed to be working together to regulate a gene called CD83, offering a clearer picture of how these cells communicate than previous methods had provided.
The most significant findings came from applying FORGE to a mouse model of Alzheimer's disease. The researchers were looking for the specific genetic programs that go awry in the brain as the disease progresses. The tool identified a distinct pattern involving a protein called Mef2c. In healthy mice, this protein helps regulate genes in the brain's support cells, known as glia. In the diseased mice, the tool found that the activity of Mef2c was significantly altered. It showed that the switches controlling Mef2c's targets were opening and closing in a way that suggested a breakdown in the normal regulatory network. The tool didn't just guess this; it found evidence across multiple layers of data. It saw changes in the gene activity, changes in the accessibility of the DNA switches, and changes in the physical binding of the protein to the DNA. This convergence of evidence from different angles gave the researchers a high degree of confidence that this specific regulatory program was indeed a key part of the disease process in these mice.
However, the researchers were careful not to overstate their findings. They knew that looking at individual cells can sometimes create an illusion of change that isn't actually there when you look at the whole animal. To be sure, they re-analyzed the data by grouping cells from the same mouse together, treating the mouse itself as the unit of measurement rather than the individual cell. When they did this, the statistical signal for the disease changes became much weaker, and many of the specific genes that looked different at the single-cell level no longer reached the threshold for statistical significance. This didn't mean the biology was wrong, but rather that the changes were subtle and the number of mice in the study was too small to prove them with absolute certainty. The researchers concluded that while the specific regulatory program involving Mef2c is a very strong candidate for further study, it remains a hypothesis that needs to be tested in larger, more powerful experiments.
The value of FORGE lies not just in the specific answers it found, but in how it found them. It provides a standardized, transparent way to move from raw data to biological insight, ensuring that the results are reproducible and that the path taken to get there is clear. By integrating different types of evidence—gene expression, DNA accessibility, and protein binding—it allows scientists to build a more complete and robust picture of cellular regulation. The tool successfully handled data from human and mouse tissues, from blood to brain to kidney, and from two different laboratory technologies, proving that it is flexible enough to be used in many different research settings. While it did not solve the mystery of Alzheimer's disease on its own, it provided a powerful new lens through which to view the complex regulatory networks that govern our cells, offering a clearer path toward understanding how these systems fail in disease.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.