spaGFM is a scalable graph foundation model for spatial transcriptomics analyses
The paper introduces spaGFM, a scalable graph foundation model that leverages random walks and self-supervised learning to generate robust, transferable representations of spatial transcriptomics data, enabling the analysis of complex tissue organization and functional consequences across diverse biological contexts.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Biology has long relied on taking a microscope to a piece of tissue and counting the cells it sees, but this approach often misses the bigger picture. A single cell does not exist in a vacuum; its behavior is dictated by the company it keeps. In a living organ, a cell's function changes depending on whether it is sitting next to a neighbor that is fighting an infection or one that is quietly repairing damage. To understand how tissues work, and why they fail in disease, scientists need to see not just the individual players, but the complex, shifting neighborhoods they form. In recent years, a new technology called spatial transcriptomics has made this possible by mapping exactly where every gene is turned on within a tissue sample, preserving the map of the neighborhood alongside the list of its residents. However, these maps are becoming so vast and detailed that they are difficult for computers to read. The data is not a simple list of words or a straight line of code; it is a tangled web of connections, where the distance and relationship between cells matter more than the cells themselves.
A team of researchers has introduced a new tool called spaGFM, designed to make sense of these intricate cellular neighborhoods. Think of it as a system that learns to read the story of a tissue by tracing the paths between its cells, rather than just looking at them in isolation. The researchers trained this system on a massive collection of data, feeding it information from 132 different tissue samples containing nearly 44 million cells. These samples came from various parts of the human and mouse body, including the lung, liver, and brain, and were captured using different high-resolution imaging machines. By studying this enormous library, the system learned to recognize the patterns of how cells organize themselves into functional groups. It does this by treating the tissue as a map of connected dots, where each dot is a cell. Instead of trying to process the whole map at once, the system takes "walks" across the map, hopping from one cell to its neighbor, and then to that neighbor's neighbor. It records these journeys as a sequence of steps, allowing it to understand the local environment of any given cell, from its immediate surroundings to the broader neighborhood it inhabits.
The power of this approach was tested in three very different biological challenges, each revealing how well the system could apply what it learned to new situations. First, the researchers asked if the system could find specific immune structures known as tertiary lymphoid structures. These are small, organized clusters of immune cells that form in tissues during chronic inflammation or cancer, and their presence often predicts how well a patient will respond to treatment. Identifying them is notoriously difficult, especially in older, lower-resolution imaging data where the boundaries between cells are blurry. The system, having been trained only on high-resolution images, was able to successfully locate these structures in lower-resolution data from lung and kidney cancer samples. It did this without needing to be retrained on the new data, simply by recognizing the spatial patterns it had learned earlier. The results showed that the system could distinguish these immune clusters from the surrounding tumor tissue more accurately than previous methods, even when the data came from a completely different type of machine.
Next, the team explored how the system understood the influence of the immune environment on cancer cells. In a study involving genetic experiments on mouse tumors, researchers had already altered specific genes in cancer cells to see how they reacted. The question was whether the reaction of a cancer cell depended on whether it was standing next to a T-cell, a type of immune cell. The system was able to look at the genetic changes in the cancer cells and correctly predict which ones were near T-cells and which were far away, simply by analyzing the patterns in their gene activity. More importantly, it could predict how a cancer cell would react to a genetic change it had never seen before, but only if that reaction depended on its proximity to an immune cell. This suggests the system had truly learned the rules of how cells talk to their neighbors, rather than just memorizing a list of gene behaviors.
Finally, the researchers applied the tool to a human kidney biopsy from a patient with diabetic kidney disease. In this condition, the tiny filtering units of the kidney, called glomeruli, become damaged and scarred. Pathologists grade this damage from mild to severe, but doing so by eye is difficult and subjective. The system was able to look at the tissue and automatically group the cells into the correct damage categories, matching the grades assigned by human experts. It even identified new areas that looked like damaged glomeruli but had been missed by the initial human review, which were later confirmed by the presence of specific molecular markers. The system also revealed that as the damage worsened, the composition of the neighborhood changed in predictable ways, with certain protective cells disappearing and others taking their place.
The significance of this work lies in its ability to scale. Previous methods for analyzing these cellular maps often struggled when the data grew too large or too complex, getting bogged down by the sheer number of connections. This new approach bypasses those limits by converting the complex web of cell-to-cell connections into a format that modern computers can process efficiently. It allows scientists to learn general rules about how tissues are built and how they break down, rules that hold true across different organs and different diseases. While the system is not a magic solution that replaces human experts, it provides a powerful new lens for viewing the microscopic world. It turns the chaotic jumble of billions of cells into a readable story of organization, offering a clearer path to understanding the fundamental principles of life and disease.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.