← Latest papers
📊 statistics

Causal Effects of Modified Treatment Policies under Positivity Violations: A Partial Identification Approach

This paper proposes a partial identification framework for estimating causal effects of modified treatment policies under positivity violations by bounding unobserved outcomes via Lipschitz continuity and introducing a novel interior-displaced projection method to ensure pathwise differentiable, asymptotically normal estimators that provide robust confidence intervals where traditional positivity-assuming methods fail.

Original authors: Taehyeon Koo, Elizabeth A. Stuart, Kara E. Rudolph, Caleb H. Miles

Published 2026-08-26
📖 5 min read🧠 Deep dive

Original authors: Taehyeon Koo, Elizabeth A. Stuart, Kara E. Rudolph, Caleb H. Miles

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine trying to understand the effect of a new policy by looking at what happened to people in the past. If the policy says, "Everyone should reduce their pesticide exposure by twenty percent," researchers can usually predict the outcome by finding people who naturally had that lower exposure level and seeing what happened to them. This works well when the data is rich and covers all the necessary ground. However, in the real world, especially when dealing with complex mixtures of chemicals or continuous variables like drug dosages, the data often has gaps. There are combinations of factors that simply do not exist in the historical records. When a policy asks for a change that lands in one of these empty zones, standard methods are forced to guess what would happen, often relying on mathematical models that stretch far beyond the available evidence. This guessing game is risky because the models might be wrong, leading to misleading conclusions about whether a policy is safe or harmful.

A team of researchers at Columbia University and Johns Hopkins University has developed a new way to handle these gaps without making dangerous guesses. Instead of trying to fill in the missing pieces with a single, rigid formula, they treat the unknown area as a zone of uncertainty. They accept that they cannot know the exact outcome for every single person the policy might affect, but they can calculate a range of possible outcomes that is guaranteed to contain the truth. Their approach splits the problem into two parts: the part where data exists and the part where it does not. For the part with data, they use standard, reliable methods to find the answer. For the part without data, they do not guess a specific number. Instead, they look at the closest person in the data who is similar to the missing case and ask a simple question: how much could the outcome possibly change if we moved just a little bit further away?

To answer this, the researchers impose a rule of smoothness. They assume that the outcome does not change wildly over short distances; if two situations are very similar, their results should be somewhat similar, too. By measuring the distance between the missing case and the closest known case, and multiplying that distance by a safety factor that represents how much the outcome might vary, they create a safe boundary. This boundary acts like a fence, ensuring that the true answer lies somewhere within the calculated range. The researchers call this a "partial identification" approach because it identifies a range of possibilities rather than a single point. They found that this method is particularly important when dealing with mixtures of chemicals, such as the various pesticides found in the environment. In these cases, the chemicals often move together in the data, making it very hard to find examples where one is low while another is high, creating large gaps in the information.

The team tested their method using computer simulations where they knew the true answer in advance. In scenarios where standard methods, which rely on guessing the missing data, failed to capture the true answer, the new method successfully kept the correct result within its calculated range. They also applied their technique to a real-world study involving pesticide exposure and maternal hypertension in a specific community. The standard methods suggested a protective effect, implying that reducing pesticide exposure would lower the risk of high blood pressure. However, when the researchers applied their new, more cautious approach, they found that this protective suggestion was not robust. Once they accounted for the uncertainty in the gaps of the data, the evidence for a protective effect disappeared. This does not mean the policy is bad, but rather that the data is not strong enough to support a confident claim that it is good.

A major hurdle the researchers had to overcome was a technical problem related to how they defined the boundary between the known and unknown data. In their initial design, the boundary was a sharp line, which made it mathematically difficult to calculate how precise their estimates were. To solve this, they introduced a clever adjustment that moved the boundary slightly inward, creating a thin, smooth layer where the transition happens. This small geometric change allowed them to use powerful statistical tools to prove that their estimates are reliable and to construct confidence intervals that are valid across different levels of uncertainty. They showed that by accepting a small amount of uncertainty in the calculation, they could gain a much clearer and more honest picture of what the data actually supports.

The work highlights a fundamental shift in how we should think about policy analysis in complex environments. Rather than forcing a single number out of incomplete data, which often leads to overconfidence, it is better to acknowledge the limits of what we know. The researchers demonstrated that by explicitly stating the assumptions about how much outcomes can change over distance, and by calculating a range that respects those assumptions, we can avoid the trap of false certainty. Their method provides a tool for scientists and policymakers to see clearly when a policy's effects are supported by evidence and when they are merely the result of a mathematical guess. In the end, the goal is not to eliminate uncertainty, which is often impossible, but to measure it accurately so that decisions can be made with a full understanding of the risks involved.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →