← Latest papers
🤖 machine learning

Generalised Transportability via Causal Abstractions

This paper proposes a model-level framework for generalized transportability grounded in causal abstraction theory, which replaces query-specific identification with a single alignment map to simultaneously transport all queries, while providing certified interval bounds for non-transportable or target-agnostic scenarios through distributionally robust optimization.

Original authors: Yorgos Felekis, Paris Giampouras, Fabio Massimo Zennaro, Theodoros Damoulas

Published 2026-08-18
📖 6 min read🧠 Deep dive

Original authors: Yorgos Felekis, Paris Giampouras, Fabio Massimo Zennaro, Theodoros Damoulas

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of cause and effect, scientists often face a frustrating dilemma: they learn how a system works in one place, but they need to know how it will behave somewhere else. Imagine a doctor who has studied how a new drug helps patients recover in a bustling city hospital, only to wonder if that same treatment will work for a rural community with different air quality and age demographics. This is the problem of transportability. For decades, the standard tools to answer this question have been like a strict gatekeeper. They look at the data and the connections between variables, and they give a binary answer: yes, the result can be moved, or no, it cannot. If the answer is yes, the tools provide a precise formula. If the answer is no, the tools simply stop talking, leaving the researcher with no information at all. This silence is a major problem because, in the real world, data is often messy, incomplete, or entirely missing from the new location.

A team of researchers at the University of Warwick and the University of Bergen has proposed a different way to think about this challenge. Instead of asking whether a single specific question can be answered, they asked whether the entire underlying model of the world can be translated. They treated the source population and the target population as two versions of the same machine, differing only in a few specific gears. By focusing on the relationship between these two machines, they developed a new framework that does not just say "yes" or "no." Instead, it provides a range of likely outcomes, a safety net that holds true even when the exact answer is impossible to calculate. Their work transforms a rigid, all-or-nothing verdict into a flexible, certified estimate that tells researchers exactly how far they can trust their conclusions.

The researchers, led by Yorgos Felekis and colleagues, built their approach on the idea that the source and target environments share a common structure. They share the same variables and the same causal connections, but some of the rules governing how those variables interact have changed. In their new method, they do not try to find a perfect formula for every single scenario. Instead, they look for a single map that can translate the behavior of the source system into the behavior of the target system. When such a perfect map exists, the translation is exact. But the true power of their work lies in what happens when a perfect map does not exist. In many real-world cases, the differences between the two populations are too complex for a single, exact translation. Here, the researchers found that they could still find the best possible approximate map. Even though this map is not perfect, the error it makes is not random. It is bounded and predictable. By calculating this error, they can draw a certified interval—a guaranteed range of values—that is certain to contain the true answer, no matter how the target population differs from the source.

To test this idea, the team created several simulated worlds and also applied their method to real-world ecological data. In one simulation, they tackled a scenario where the standard tools of the field would have declared the problem unsolvable because of hidden, unmeasured factors. In this "unsolvable" case, their new method successfully produced a narrow, informative range that contained the true answer. In another test, they simulated a situation where no data from the target location existed at all. Even without seeing the target, their method learned a map from the source data that remained reliable across a wide variety of possible target environments. They also tested their framework on a real dataset concerning river ecosystems, specifically looking at how tree cover affects oxygen levels in water. In this real-world case, they found that their method could identify the correct range of effects even when the target environment was significantly different from the source.

A key feature of their approach is the ability to use prior knowledge to sharpen the results. If a researcher knows that a specific change is likely to happen in a certain direction—for example, if they know a treatment effect is likely to get stronger rather than weaker—they can encode this belief into the model. The researchers showed that using this directional knowledge makes the certified intervals much tighter and more useful. Without this knowledge, the intervals are wider to be safe, but they still hold true. When the researchers compared their method to the old way of doing things, they found that the new method provided useful information in situations where the old method gave up completely. In some tests, the new method's intervals were several times tighter than the general bounds, offering a much clearer picture of the truth.

The researchers also explored how to choose the right settings for their model when no target data is available. They found that in some cases, looking at the observational data from the target location could help select the right level of caution for the model. However, they discovered that this strategy does not always work. In cases where the difference between the two populations was purely environmental, the standard way of checking the data suggested using a very small safety margin, which would have led to a wrong answer. Their analysis showed that relying solely on observational data to set these limits can be risky. Instead, they argued that the safety margin should be set based on a clear understanding of how much the environment might change, ensuring the final guarantee remains valid even if the data looks deceptively similar.

Ultimately, this work changes the conversation about moving knowledge from one place to another. It moves the field away from a binary view where a problem is either solvable or impossible. Instead, it offers a continuum where every problem has a solution, provided one is willing to accept a range of possibilities rather than a single number. The researchers demonstrated that by treating transportability as a problem of structural alignment between two models, they could provide certified guarantees even in the most difficult scenarios. Their method does not just tell scientists what they can know; it tells them how sure they can be, turning the silence of the old methods into a clear, actionable voice.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →