← Latest papers
📈 economics

Prediction Sets and Conformal Inference with Interval Outcomes

This paper develops consistent and finite-sample valid methods for constructing optimal prediction sets for censored or interval-valued outcomes by characterizing oracle prediction intervals under partial identification and applying tailored conformal inference to account for irreducible, partial identification, and sampling uncertainties.

Original authors: Weiguang Liu, Áureo de Paula, Elie Tamer

Published 2026-02-24
📖 5 min read🧠 Deep dive

Original authors: Weiguang Liu, Áureo de Paula, Elie Tamer

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to guess the price of a house in a new city. In a perfect world, you would know the exact price of every house sold there. But in the real world, data is messy. Sometimes, people don't say, "This house sold for $500,000." Instead, they say, "It sold for somewhere between $450,000 and $550,000." Or perhaps they only say, "It sold for more than $500,000."

This is what economists call interval data. It's like looking at a map where some cities are marked with a precise dot, but others are marked with a fuzzy cloud.

This paper by Liu, de Paula, and Tamer is about how to make the best possible guess (a prediction) when your data is fuzzy, and how to be confident that your guess is right.

Here is the breakdown of their work using simple analogies:

1. The Goal: The "Perfect Net" (The Oracle)

Imagine you are fishing. You want to catch a fish (the true outcome), but you don't know exactly where it is.

  • The Point Guess: Most people try to guess the exact spot where the fish is. "It's at coordinate X." If you miss by an inch, you fail.
  • The Prediction Set: This paper suggests casting a net. You want a net that is big enough to catch the fish 95% of the time (so you don't miss), but small enough to be useful. If your net covers the whole ocean, you technically "caught" the fish, but you learned nothing.

The authors call the "perfect net" the Oracle Prediction Set. It is the smallest possible net that still guarantees you catch the fish 95% of the time.

2. The Problem: The "Fuzzy" Data

Usually, to build this perfect net, you need to know exactly where the fish are swimming. But here, the data is censored or interval-valued.

  • The Analogy: Imagine you are trying to draw a map of where fish swim, but your sonar only tells you, "There is a fish somewhere in this 10-foot square," or "There is a fish somewhere above this line." You don't know the exact spot.
  • The Challenge: If you try to use old methods (which assume you know the exact spot), your net will either be too loose (wasting space) or too tight (missing the fish). The authors had to invent a new way to draw the net using only these fuzzy squares.

3. The Solution: Two-Step Magic

Step A: The "Smart Estimator" (Consistency)

First, the authors built a mathematical machine that looks at all those fuzzy squares and figures out the shape of the "perfect net."

  • How it works: They use a technique called kernel smoothing. Imagine you have a pile of sand (your data points). You pour a little water on it, and the sand settles into a smooth hill. This hill represents the probability of where the fish are.
  • The Twist: Because the data is fuzzy, the "hill" might have two peaks (bimodal). Maybe fish like to hang out in deep water and shallow water, but not in the middle. A simple method might draw one big net covering both. The authors' method is smart enough to draw two separate nets (one for deep, one for shallow) if that's where the fish actually are. This makes the net much smaller and more efficient.

Step B: The "Safety Net" (Conformal Inference)

Even with a smart machine, you might make a mistake because you only have a limited amount of data (sampling error). You need a guarantee that your net works right now, not just in the long run.

  • The Analogy: Think of Conformal Inference as a "calibration test."
    1. You build your net using 75% of your data (the training set).
    2. You take the remaining 25% (the calibration set) and throw them into your net.
    3. You ask: "How many of these test fish did my net catch?"
    4. If the net was too small and missed some, you widen the net just enough to catch them. If the net was huge and caught everything easily, you shrink it slightly.
  • The Result: This guarantees that no matter what the data looks like, your final net will catch the fish at least 95% of the time. It's like a "money-back guarantee" for your prediction.

4. Real-World Applications

The authors tested their method on two real-world scenarios:

  • UK Job Salaries: Many job ads don't list a specific salary; they say "50k50k–70k." The authors used their method to predict salary ranges across different UK cities. They found that in London, the "net" was huge (salaries vary wildly from low to very high), while in smaller towns, the net was tighter. This gives a much clearer picture of the job market than just guessing an "average" salary.
  • US Income Data: In surveys, people often refuse to give their exact income but will say, "I make between $40k and $60k." Traditional methods often throw this data away or guess a single number (imputation), which can be wrong. The authors' method keeps the "fuzzy" range and builds a prediction net around it, showing that people with more education have a wider range of possible incomes (a bigger net), not just a higher average.

Summary

This paper is about making better guesses when you don't have perfect information.

  1. Don't guess a single number; guess a range (a net).
  2. Make the net as small as possible while still being safe (efficient).
  3. Handle fuzzy data (intervals) without throwing it away or guessing the exact numbers.
  4. Use a calibration test (conformal inference) to guarantee your net works, even with small sample sizes.

It's a toolkit for turning "I don't know exactly" into "I know it's definitely in this specific zone."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →