Pretrained gene representations transfer mean expression more broadly than spatial patterns in virtual spatial transcriptomics
This study demonstrates that in virtual spatial transcriptomics, pretrained gene representations primarily enhance the prediction of a gene's mean expression across tissues rather than its spatial variation, revealing that cross-gene generalization is driven more by broad expression transfer than by the recovery of specific spatial patterns.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine a city where every building holds a secret library, and inside each library, thousands of books are open to different pages. In the human body, these buildings are cells, and the books are genes. For decades, scientists have been able to read which books are open in a single cell, but they have struggled to see where those cells are located within the complex tissue of an organ. To get a full map of gene activity, researchers traditionally need to take a physical slice of tissue and run expensive, time-consuming tests on it. This limits how much of the body they can study at once. A newer approach, called virtual spatial transcriptomics, tries to solve this by using artificial intelligence. Instead of running a new test on every gene, the computer looks at a standard photograph of the tissue and predicts which genes are active and where. The hope is that by teaching the computer to recognize patterns in the tissue image, it can guess the location of genes it has never seen before, effectively extending the map to new areas without new lab work.
The core question researchers asked was whether these computer models are truly learning the complex, shifting patterns of gene activity across a tissue, or if they are simply getting very good at guessing the average amount of a gene present. When a model predicts a gene's location, it is doing two things at once: estimating the overall level of that gene in the tissue, and figuring out how that level changes from one spot to another. A model could appear successful by simply assigning the same average value to every location, which would look like a good prediction of the total amount but would fail to capture the intricate details of where the gene is actually concentrated. To find the truth, the researchers built a system that could separate these two abilities. They tested the models on genes that the computer had never seen during its training, asking if the model could still predict the gene's average level and its specific spatial pattern in new patients.
The researchers tested their ideas on four different groups of human tissue samples, including three distinct regions of the brain and a type of breast cancer. They trained their models on a large set of genes and then challenged them to predict the behavior of hundreds of genes that were completely new to the system. They used two different types of pre-trained gene representations, which are essentially digital summaries of a gene's identity derived from vast amounts of biological data. One type of summary comes from the DNA sequence surrounding the gene, while the other comes from how the gene behaves in single cells. When the models tried to predict the new genes, the results showed a clear split in performance. The models were remarkably successful at predicting the average expression level of these unseen genes. In fact, when the models improved their overall accuracy, more than 90 percent of that improvement came from simply getting the average amount of the gene right.
However, the ability to recover the specific spatial patterns—the "where" of the gene activity—was much more limited. The researchers found that the models only improved their spatial predictions for genes that already showed high variation in the training data. For most genes, the computer could tell you how much of the gene was present on average, but it could not reliably tell you where in the tissue that gene was concentrated. To prove that the models were not actually using the tissue images to find these patterns, the researchers ran a second test. They built a simplified version of the model that looked only at the gene's digital summary and ignored the tissue photograph entirely. Surprisingly, this image-free model retained nearly all of the improvement in predicting the average gene levels that the full model had achieved. This demonstrated that the pre-trained gene summaries were carrying the information about the gene's average level, but the tissue images were not adding much new information about the gene's specific location for these unseen targets.
The study also revealed that the type of gene representation mattered for the spatial patterns. While both types of summaries helped predict average levels, they differed in which specific genes showed improved spatial accuracy. Some genes benefited from one type of summary, while others benefited from the other. This suggests that the ability to transfer knowledge from known genes to new ones is not a single, uniform skill. Instead, it is a combination of two distinct capabilities: a broad ability to estimate how much of a gene is present, and a selective, difficult ability to map its precise location. The researchers concluded that while these virtual tools are powerful for expanding the list of genes we can study, they are not yet fully capturing the complex spatial architecture of the tissue for every gene. The success of these models in predicting average levels is a significant step forward, but the challenge of accurately reconstructing the detailed spatial patterns of new genes remains a work in progress.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.