Partial Identification under Causal Orders by Linear Programming
This paper proposes a linear programming framework for the partial identification of counterfactual and nested counterfactual queries by leveraging structural orderings inherent to the queries themselves, thereby eliminating the need for a fully specified causal graph while providing tight, data-compatible bounds.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of cause and effect, we often find ourselves asking "what if" questions that go beyond simple observation. We want to know not just what happened, but what would have happened if a different choice had been made. Did a specific treatment save a patient's life, or would they have recovered anyway? Did a policy change improve an economy, or was the improvement inevitable? To answer these questions with certainty, scientists usually need a complete map of the forces at play, a detailed diagram showing exactly how every factor influences every other. This map is called a causal graph. Without it, the answers remain hidden in a fog of possibilities. For decades, researchers have struggled to make sense of these "what if" scenarios when that map is missing or incomplete, often forced to rely on rough guesses or to admit that the question cannot be answered at all.
Two researchers, Eric Rossetto and Alessandro Antonucci, have developed a new way to navigate this fog. They realized that even without a full map, the very nature of a "what if" question contains its own internal logic. If you ask how changing one thing affects another, the question itself implies a sequence: the change must happen before the result. By focusing on this inherent order, the researchers created a method to calculate the tightest possible range of answers for these counterfactual questions, without ever needing to assume a specific causal map. Their work transforms a complex, often impossible guessing game into a structured calculation that yields precise boundaries for what is possible, given only the data we have and the logic of the question itself.
The core of their discovery lies in recognizing that every inquiry about cause and effect carries a hidden instruction about time and sequence. When we ask if a treatment caused an outcome, we are implicitly saying the treatment came first. The researchers showed that this simple ordering is enough to build a mathematical framework that can test every conceivable way the world could be arranged, consistent with that order. Instead of trying to guess the exact shape of the causal web, they built a system that explores the entire landscape of possibilities allowed by the data and the question's logic. They proved that this system can find the absolute best and worst-case scenarios for any such question, providing a range that is guaranteed to contain the true answer.
To achieve this, the team turned to a powerful mathematical tool known as linear programming, which is essentially a method for finding the best outcome in a system with many constraints. They translated the problem of guessing causal effects into a set of rules that a computer could solve. Imagine a vast space of all possible stories about how the world works. Some stories fit the data we have collected; others do not. The researchers' method filters this space, keeping only the stories that respect the order implied by the question and the observed data. Within this filtered space, they calculated the lowest and highest possible probabilities for the event in question. They demonstrated that these calculated limits are not just rough estimates, but the sharpest possible boundaries. They proved that there are real, concrete models of the world that fit all the data and the question's logic, and that produce results exactly at these upper and lower limits. This means the range they provide is not an artifact of their method, but a true reflection of what is possible given our current knowledge.
The power of this approach becomes clear when looking at questions that have stumped experts before. The researchers tested their method on several classic problems from scientific literature, including scenarios involving medical treatments and social policies. In one famous case involving university admissions, where researchers have long debated whether gender influenced acceptance rates, the new method provided a wide but honest range of possibilities. Previous analyses that assumed a specific, rigid structure for how gender and department choices interacted produced a single, precise number suggesting no discrimination. However, the researchers' method, which refused to make those unverified structural assumptions, showed that the true answer could be anywhere within a much broader interval. This wider range did not mean the answer was unknown in a useless way; rather, it revealed that the precise number from earlier studies relied heavily on assumptions that might not be true. By removing those assumptions, the researchers showed that the conclusion of "no discrimination" was far less certain than previously thought.
This work also extends to more complex questions involving "nested" scenarios, where one "what if" is buried inside another. For example, asking what would happen if a patient took a drug, but only if they had also been assigned to a specific group in a trial. These layered questions are notoriously difficult to answer without a full causal map. The researchers showed that their method handles these nested layers just as effectively as simple ones. They demonstrated that by breaking down these complex queries into their logical components, the same mathematical machinery could find the tightest possible bounds. In their tests, they found that trying to combine separate, simpler answers to build a complex one often led to much wider, less useful ranges. Their direct approach to the complex question, however, yielded significantly tighter and more informative results.
The researchers were careful to acknowledge the trade-off involved in their method. Because they do not assume a specific causal map, their calculated ranges are naturally wider than those derived from studies that make strong assumptions about how variables are connected. A study that assumes a specific diagram might produce a very narrow, precise answer, but that precision comes at the cost of potentially being wrong if the assumed diagram is incorrect. The new method sacrifices that narrowness for reliability, ensuring that the true answer is never excluded from the range. They proved that this reliability is not a weakness but a strength, especially in fields like medicine and social science where the true causal structures are often unknown or debated.
Ultimately, this research offers a new way to think about uncertainty. It suggests that we do not need to know everything about the world to make meaningful progress in understanding cause and effect. By leveraging the logical structure inherent in our questions, we can extract the maximum amount of information from our data without overstepping into unsupported assumptions. The researchers provided a tool that allows scientists to say, with mathematical certainty, "The answer lies somewhere between X and Y," even when they cannot say exactly where. This ability to define the boundaries of the possible, without needing a complete map of the territory, represents a significant step forward in our ability to reason about the past and the future in a world where complete knowledge is rarely available.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.