← Latest papers
📊 statistics

Testing the Missing Completely at Random Assumption for Functional Data

This paper introduces a novel statistical testing framework, utilizing both deterministic partitions and clustering-based approaches, to validate the Missing Completely at Random (MCAR) assumption for functional data observed only on subsets of their domain.

Original authors: Maximilian Ofner, Siegfried Hörmann, David Kraus, Dominik Liebl

Published 2026-07-07
📖 5 min read🧠 Deep dive

Original authors: Maximilian Ofner, Siegfried Hörmann, David Kraus, Dominik Liebl

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery about a collection of stories. These stories are "functional data"—think of them as continuous lines or curves, like a heart rate monitor tracing a line on a screen, or a graph showing electricity prices over a week.

Usually, these stories are told from beginning to end. But in the real world, sometimes the recorder breaks, the battery dies, or the sensor gets confused. This means you only have fragments of the stories. Some are complete; others are cut off early or start late.

The big question for statisticians is: Why are these stories missing pieces?

The Core Mystery: "Missing Completely at Random" (MCAR)

The paper focuses on a specific assumption called Missing Completely at Random (MCAR).

  • The Ideal Scenario (MCAR): Imagine the recorder breaks because of a random coin flip. The story of a heart rate doesn't matter; the machine just happens to fail. In this case, the missing parts are just as interesting as the present parts. If you look at the people whose stories are complete, they are just a random slice of the whole group.
  • The Problem: If the recorder breaks because the heart rate got too high (or the electricity price got too low), then the missing data is not random. The missingness is linked to the data itself. If you ignore this, your analysis is like trying to guess the average height of a basketball team by only measuring the people who didn't get injured. You'll get the wrong answer.

For a long time, statisticians had no good way to check if the "missingness" was truly random or if it was sneaking in a bias. This paper builds a new toolkit to solve that.

The Detective's Toolkit: Splitting the Clues

The authors propose a clever way to test if the missingness is random. They use a simple two-step process:

  1. Sort the Clues: They take all the "observation patterns" (the shapes of the missing pieces) and split them into two groups.
    • Simple way: Group 1 has "Complete Stories," and Group 2 has "Broken Stories."
    • Smart way: They use a computer algorithm (clustering) to find natural groups. For example, maybe one group has stories missing from the beginning, and another has stories missing from the end.
  2. Compare the Stories: Once sorted, they ask: "Do the stories in Group 1 look different from the stories in Group 2?"
    • If the missingness is random (MCAR), the two groups should look exactly the same. The fact that one group has missing pieces shouldn't change the shape of the story.
    • If the missingness is not random, the two groups will look different. For instance, if high electricity prices cause the recorder to fail, the "Broken" group will have lower prices than the "Complete" group.

The Two Main Tests

The paper introduces two ways to compare these groups, like using two different magnifying glasses:

  • The Average Test (Mean Comparison): This looks at the "average shape" of the curves in both groups. If the average heart rate of the "Broken" group is significantly different from the "Complete" group, the test raises an alarm.
    • Analogy: Imagine comparing the average height of two groups of people. If one group is noticeably taller, something is wrong with how they were selected.
  • The Full Picture Test (Distribution Comparison): This is a more powerful, all-seeing eye. It doesn't just look at the average; it looks at the entire shape and spread of the data. It checks if the whole story is different, not just the average.
    • Analogy: The average test might miss it if the "Broken" group has some very tall people and some very short people that balance out to the same average height. The Full Picture test sees that the variety of heights is different and catches the problem.

Real-Life Cases: What Did They Find?

The authors tested their new toolkit on three real-world datasets:

  1. Heart Rates (The "All Clear"): They looked at heart rate data from 878 people. Some devices failed, but the test said, "No problem!" The missing data looked random. The "Broken" group and the "Complete" group had the same heart rate patterns. This confirmed what doctors suspected: the devices just failed randomly.
  2. Electricity Prices (The "Smoking Gun"): They looked at electricity prices. Here, the test screamed, "Something is wrong!" The data showed that when prices were high, the data often went missing. The "Broken" group had systematically different prices than the "Complete" group. This makes sense: in the electricity market, high demand (and high prices) might cause data transmission issues. The test successfully caught this bias.
  3. Temperature Sensors (The "Broken Sensor"): They looked at temperature data from Graz, Austria. Some days had missing data in the second half of the day. The test found a significant difference, suggesting the sensor might be failing specifically when it gets too hot (or for some other technical reason). This helps engineers know exactly where to look for the broken part.

Why This Matters

Before this paper, statisticians had to guess if their missing data was safe to use. If they guessed wrong, their conclusions could be completely wrong.

This paper gives them a statistical lie detector. It allows researchers to:

  • Check if their data is biased before they start analyzing it.
  • Use a smart computer method (clustering) to find the best way to split the data for testing.
  • Use visual tools (graphs with confidence bands) to see exactly where the data is behaving differently.

In short, the authors built a bridge between the messy reality of broken data and the need for clean, reliable statistics, ensuring that when we analyze functional data, we aren't just analyzing the "easy" parts while ignoring the "hard" ones.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →