← Latest papers
📊 statistics

Counterfactual Forecasting for Panel Data

This paper introduces FOCUS, a novel method for forecasting counterfactual outcomes in panel data with missing entries by extending matrix completion to leverage time series dynamics of latent factors, thereby achieving superior prediction accuracy and establishing theoretical guarantees for applications like mobile health studies.

Original authors: Navonil Deb, Raaz Dwivedi, Sumanta Basu

Published 2026-02-03
📖 5 min read🧠 Deep dive

Original authors: Navonil Deb, Raaz Dwivedi, Sumanta Basu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a coach trying to predict how your team will perform in next week's games. You have a spreadsheet with data on every player's past performance, but there are two big problems:

  1. Missing Data: Some players missed practice or didn't play in certain games, leaving blank spots in your spreadsheet.
  2. Hidden Patterns: The players aren't just random individuals; they are influenced by invisible "moods" or "team dynamics" (like fatigue, morale, or weather) that change over time and affect everyone differently.

The paper introduces a new tool called FOCUS (Forecasting Counterfactuals Under Stochastic Dynamics) to solve this. Here is how it works, broken down into simple concepts:

The Problem: The "What If" Guessing Game

In science and policy, we often want to know: "What would have happened if we had done X instead of Y?" This is called a counterfactual.

For example, in a mobile health app, we might want to know: "If we had sent a walking reminder to this user at 5 PM, how many steps would they have taken?" But we can't send the reminder and not send it at the same time. We only see what happens when we do send it. The "what if" scenario (the counterfactual) is missing.

Usually, statisticians try to fill in these missing blanks by looking for patterns in the data. However, most existing tools treat time like a static photo—they look at the past to guess the future but ignore the fact that time flows. They miss the fact that if a user is tired today, they are likely to be tired tomorrow.

The Solution: FOCUS

The authors built FOCUS to act like a time-traveling detective. Instead of just looking at the data as a flat grid, FOCUS understands that the hidden "moods" (called latent factors) move and evolve like a story.

Here is the step-by-step process FOCUS uses:

  1. Finding the Invisible Threads (Factor Estimation):
    Imagine the data is a giant tapestry with holes in it. FOCUS first looks at the visible threads to figure out the underlying pattern of the tapestry. It uses a mathematical technique (similar to finding the main themes in a song) to estimate these hidden "moods" even when data is missing.

  2. Predicting the Next Chapter (Time Series Dynamics):
    This is the magic ingredient. Once FOCUS knows what the "moods" were yesterday and the day before, it doesn't just guess randomly for tomorrow. It uses a time machine (specifically, a mathematical model called VAR) to predict how these moods will evolve.

    • Analogy: If you know a car is moving north at 60 mph, you can predict where it will be in an hour. If you just look at a photo of the car, you can't. FOCUS looks at the car's speed and direction (the time dynamics) to predict the future.
  3. Filling in the Blanks (Forecasting):
    With the predicted "moods" for the future, FOCUS fills in the missing spots in the spreadsheet. It tells you, "Based on the hidden patterns and how they usually move, here is what the user's step count would have been if we had nudged them."

Why is this better than the old ways?

The paper compares FOCUS to two other popular methods:

  • mSSA: This method is like a painter who only looks at the colors on the canvas but ignores the brushstrokes. It assumes the patterns are rigid and don't change randomly.
  • SyNBEATS: This is like a very heavy, slow robot that needs a perfect, complete photo of the past to work. If there are holes in the data, it struggles.

FOCUS wins because:

  • It handles missing pieces: It can work even when the data is full of holes (which happens often in real life).
  • It understands randomness: It knows that the future isn't just a perfect copy of the past; it accounts for the "noise" and random changes in the hidden moods.
  • It's fast: It runs much quicker than the heavy robot (SyNBEATS), making it practical for large datasets.

Real-World Test: The HeartSteps Study

The authors tested FOCUS on a real mobile health study called HeartSteps. In this study, users received prompts to walk.

  • The Finding: The researchers noticed that a user's activity in one time slot (e.g., 4 PM) was strongly linked to their activity in the next slot (e.g., 5 PM). If they walked a lot at 4 PM, they were likely to walk less at 5 PM (perhaps because they were tired).
  • The Result: Because FOCUS understood this "tiredness" pattern (the time dynamics), it predicted future step counts much more accurately than the other methods. It successfully used the "story" of the user's day to guess the next chapter.

The Bottom Line

FOCUS is a new way to predict the future in situations where data is messy and incomplete. By treating hidden patterns as moving, evolving stories rather than static pictures, it gives us a much clearer view of "what would have happened" if we had made different choices. This helps researchers and policymakers make better decisions based on more accurate predictions.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →