← Latest papers
📊 statistics

Partial Identification under High-Dimensional Potential Outcomes and Confounders via Optimal Transport

This paper proposes a novel estimator for partial identification in high-dimensional causal settings that overcomes the curse of dimensionality by decomposing the optimal transport problem into a low-dimensional signal subspace and a high-dimensional residual subspace, where the latter is efficiently recovered using the Sliced Wasserstein distance to yield tighter, more informative causal bounds.

Original authors: Yunfeng Wang, Zhiheng Zhang, Zijun Gao

Published 2026-06-02
📖 5 min read🧠 Deep dive

Original authors: Yunfeng Wang, Zhiheng Zhang, Zijun Gao

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: The "Missing Puzzle Piece" Problem

Imagine you are a detective trying to figure out the true effect of a new medicine. You have data on patients who took the medicine and those who didn't. However, there's a catch: for every patient, you can only see one outcome (what happened after they took the medicine, or what happened if they didn't), but never both at the same time.

In statistics, this is called Partial Identification. You can't pinpoint the exact answer (the "point identification"), but you can draw a box around the possible answers. The goal is to make that box as small and tight as possible so your decision is more precise.

To make this box tighter, you look at other factors (like age, weight, or blood pressure) called confounders. If you can group patients who are very similar, you can get a better estimate.

The Problem: The "High-Dimensional" Trap

The paper argues that modern data is often high-dimensional. Imagine instead of just looking at age and weight, you are looking at 30 different health metrics, or even thousands of genetic markers.

When you try to compare patients across 30 or 100 dimensions, standard mathematical tools break down. It's like trying to find the shortest path between two cities on a map that has 1,000 dimensions instead of 2. The math gets so messy and computationally heavy that the results become useless or wildly inaccurate. This is known as the "curse of dimensionality."

Existing methods try to fix this by ignoring most of the data. They say, "Let's just look at the top 5 most important factors and throw the rest away."

  • The Analogy: Imagine you are trying to guess the total weight of a suitcase. You weigh the heavy books (the "signal") but decide to ignore the clothes, socks, and toiletries (the "residual") because there are too many of them. You get a number, but it's definitely too low because you threw away a lot of weight.

The Solution: The "Signal and Slicing" Method (CSS)

The authors propose a new method called Conditioned Subspace–Slicing (CSS). Instead of throwing away the "leftover" data, they find a clever way to weigh it without doing the impossible math.

Here is how it works, broken down into two steps:

1. The Signal Subspace (The Heavy Books)

First, the method identifies the few dimensions where the biggest differences between the two groups (medicine vs. no medicine) actually happen.

  • Analogy: You find the heavy books in the suitcase and weigh them precisely. This is the "Signal."

2. The Sliced Residual (The Clothes)

This is the paper's innovation. Instead of ignoring the clothes (the "Residual"), they use a technique called Sliced Wasserstein Distance.

  • The Analogy: Imagine you can't weigh the whole pile of clothes at once. So, you take a knife and slice the suitcase into thin, one-dimensional strips (like slicing a loaf of bread). You weigh each slice individually.
  • Why this works: Weighing a single slice is easy and fast, even if the suitcase is huge. By averaging the weights of all these random slices, you get a very good estimate of the total weight of the clothes.
  • The Result: You add the precise weight of the books (Signal) to the estimated weight of the clothes (Residual).

Why This Is Better

The paper proves mathematically that this new method is:

  1. Valid: It never overestimates the effect. It always provides a "safe lower bound" (a conservative estimate that is guaranteed to be true).
  2. Tighter: Because it recovers the weight of the "clothes" (the residual data) that other methods throw away, the final box of possible answers is much smaller and more informative.
  3. Efficient: It doesn't require supercomputers. The "slicing" trick makes the math fast enough to run on high-dimensional data.

The "Spiked" Assumption

The method works best under a specific condition the authors call the Extended Spiked Transport Model.

  • Analogy: Imagine the suitcase is mostly empty space, but it has a few very dense, heavy objects (the signal) and a lot of fluffy, evenly distributed cotton (the residual).
  • If the "fluffy cotton" is spread out evenly (isotropic), the "slicing" method works perfectly to estimate its weight. The paper shows that in many real-world scenarios, the "noise" or leftover data behaves like this fluffy cotton, making the method very effective.

Summary of Results

The authors tested this on:

  1. Fake Data: They created a scenario where they knew the exact answer. The new method (CSS) got much closer to the true answer than the old methods (which just looked at the top 5 factors).
  2. Real Data: They used a medical dataset about heart catheterization. Even without knowing the "true" answer, the new method produced a tighter, more useful range of possibilities than the old methods.

In short: The paper gives us a new tool to solve complex causal questions in high-dimensional data. Instead of ignoring the messy, complex parts of the data, it uses a "slicing" trick to estimate them efficiently, resulting in much sharper and more reliable answers.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →