← Latest papers
🧬 biology

On the Recoverability of Causal Relations from Bulk Gene Expression Data

This paper demonstrates that recovering causal relations from aggregated bulk gene expression data is theoretically possible only under strict linear aggregation and affine structural equation assumptions, which are empirically unsupported by real-world datasets, thereby cautioning against such causal inference without strong additional constraints.

Original authors: Gongxu Luo, Boyang Sun, Kun Zhang

Published 2026-06-02
📖 4 min read☕ Coffee break read

Original authors: Gongxu Luo, Boyang Sun, Kun Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are trying to understand how a complex machine works by looking at a single, blurry photograph of the whole thing, rather than seeing the individual gears and springs moving inside. This is essentially the challenge scientists face when studying bulk gene expression data.

Here is a simple breakdown of what this paper discovered, using everyday analogies.

The Problem: The "Smoothie" Effect

In biology, scientists often want to know how genes talk to each other. Does Gene A turn Gene B on? Does Gene C stop Gene D?

To find this out, researchers usually take a sample of tissue (like a drop of blood) and measure the RNA inside. However, a drop of blood contains millions of individual cells. Traditional methods don't look at each cell one by one; instead, they blend all the cells together into a "smoothie" (this is called bulk expression).

  • The Analogy: Imagine you have a room full of people (cells) shouting different instructions. If you record the room with a single microphone (bulk data), you only hear a loud, mixed-up roar. You lose the specific details of who said what to whom.
  • The Hope: For years, scientists have tried to use computer algorithms to figure out the original "instructions" (causal relationships) just by listening to that mixed-up roar. They hoped that if they analyzed the noise carefully, they could reconstruct the original conversation.

The Big Discovery: The Smoothie Hides the Truth

This paper asks a fundamental question: Is it actually possible to reconstruct the original conversation just from the mixed-up roar?

The authors say: Generally, no.

They proved mathematically that when you blend cells together, you often destroy the very clues needed to figure out cause-and-effect. The "smoothie" changes the rules of the game.

The Two Rules for Recovery

The researchers found that you can only recover the original causal relationships if two very strict conditions are met. Think of these as the "Golden Rules" for un-blending the smoothie:

  1. The "Straight Line" Rule (Functional Consistency):
    The way genes affect each other must be perfectly straight and simple (linear).

    • Analogy: Imagine Gene A adds exactly 10 units of energy to Gene B, no matter what. If Gene A adds 10, and Gene B adds 10 to Gene C, the math is simple.
    • The Reality: In real life, genes are messy. Sometimes Gene A adds 10, sometimes 50, sometimes it stops working entirely. It's a curve, not a straight line. The paper shows that if the relationship is curved (non-linear), blending the cells destroys the pattern forever.
  2. The "Perfect Blend" Rule (Conditional Independence):
    The way the cells are mixed together must be a simple average or sum.

    • Analogy: If you mix red and blue paint to get purple, you can mathematically reverse it only if you mixed them in a perfectly predictable way. If the mixing process is weird (like taking the "middle" color or the "brightest" color), you can't reverse it.
    • The Reality: The paper proves that only simple averaging works. If the biological system is complex, the math breaks.

The "Real World" Test

The authors didn't just do math on paper; they tested this with real data.

  • They looked at synthetic data (computer-generated scenarios where they knew the answer). As predicted, when the genes acted in simple, straight-line ways, the math worked. When they acted in complex, curved ways, the math failed completely.
  • They looked at real human gene data (both from bulk samples and single cells). They found that genes do not act in simple, straight lines. The relationships are complex and curved.

The Conclusion

Because real genes behave in complex, non-linear ways, and because bulk data blends cells together, you cannot reliably figure out the true cause-and-effect relationships between genes just by looking at bulk data.

  • The Metaphor: Trying to find the specific conversation between two people in a crowded room just by listening to the average noise level is impossible. The noise (aggregation) has erased the specific details.

What This Means for Science

The paper warns scientists: Don't trust causal maps made from bulk data unless you have strong extra proof.

If a study claims "Gene A causes Gene B" based only on bulk data, that claim might be an illusion created by the blending process, not a real biological fact. To get the truth, scientists likely need to look at individual cells (single-cell data) or make very strong, specific assumptions that the paper suggests are rarely true in nature.

In short: The "smoothie" of bulk data is too messy to let us see the individual ingredients' recipes. We need to look at the ingredients separately to understand the recipe.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →