← Latest papers
📊 statistics

Two-stage imputation of longitudinal anthropometric data with cross-reference harmonisation: a simulation study

This simulation study evaluates a reproducible, two-stage method for imputing missing longitudinal weight and height data that combines within-subject linear interpolation with growth reference-based estimation, demonstrating high accuracy and the ability to explicitly handle and audit differing reference standards across data sources.

Original authors: Flavia Alves

Published 2026-06-10
📖 5 min read🧠 Deep dive

Original authors: Flavia Alves

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to reconstruct a movie of a child's growth, but the film reel is damaged. Some frames are missing (missing data), and worse, some scenes were filmed using two different types of cameras that record color slightly differently (different growth standards like WHO vs. CDC). If you just guess the missing parts or ignore the camera differences, your final movie will be blurry or distorted.

This paper introduces a simple, two-step "film restoration" method to fix these missing growth measurements (weight and height) while keeping a clear record of exactly how every single piece of the puzzle was found.

Here is how the method works, using everyday analogies:

The Problem: Missing Frames and Mixed Cameras

In health research, we track people's weight and height over time. Often, people miss appointments, leaving gaps in the data. Also, some studies use the WHO growth charts (like a "European camera") and others use CDC charts (like an "American camera"). Even if a child is the same size, these two charts might say they are at a different "percentile" (a ranking compared to other kids). If you mix these datasets without noticing, you introduce a hidden error, like trying to edit a movie where the lighting changes randomly from scene to scene.

The Solution: A Two-Stage Restoration Process

The authors propose a "Two-Stage" method to fill in the blanks, prioritizing what we know about the specific person before guessing based on the general population.

Stage 1: The "Connect-the-Dots" Method (Within-Subject Interpolation)

  • The Analogy: Imagine you have a photo of a child at age 5 and another at age 7, but the photo for age 6 is missing. The easiest, most logical guess is to draw a straight line between the age 5 and age 7 photos. You assume the child grew steadily in between.
  • What the paper does: If a person has measurements before and after a missing visit, the computer simply draws a straight line between them to fill the gap. It uses only that specific person's data. This is the most accurate way to fill a gap because it respects that specific child's growth pattern.

Stage 2: The "Growth Chart" Method (LMS Centile Imputation)

  • The Analogy: What if you have no photos at all for a certain year, or the child only has one photo? You can't draw a line. Instead, you look at a standard "growth chart" (like the WHO or CDC charts). You find where that child usually sits on the chart (e.g., the 50th percentile, meaning they are average size), and you assume they stayed on that same "track" until the next time you saw them.
  • What the paper does: If Stage 1 can't fill a gap, the method looks at the child's known measurements to figure out their "growth track" (centile). It then assumes the child stayed on that track and calculates what their weight/height should be at the missing time, using the specific growth chart (WHO or CDC) that the original data source used.
  • The "Explicit Parameter" Twist: Crucially, this method forces the researcher to say, "I am using the WHO chart for Source A and the CDC chart for Source B." It doesn't hide this choice. It records exactly which "camera" was used for every single guess.

The "Provenance" Label: The Receipt for Every Guess

The paper emphasizes that every number in the final dataset gets a "label" or a receipt:

  1. Observed: We actually measured this.
  2. Interpolated: We connected the dots (Stage 1).
  3. Growth Reference: We guessed based on the chart (Stage 2).

This is vital because if a researcher wants to be extra careful, they can say, "I will only trust the 'Observed' and 'Interpolated' numbers, and ignore the 'Growth Reference' guesses," or they can analyze them separately to see if the "guesses" change the results.

The Results: How Good is the Restoration?

The authors tested this method using synthetic data (computer-generated fake data that looks like real growth studies). They deliberately hid 20% of the known numbers and tried to guess them back.

  • The Outcome: The method successfully filled 100% of the gaps.
  • The Accuracy: The guesses were very close to the truth. For weight, the average error was about 1.78 kg (roughly 3.5% off). For height, it was about 2.84 cm (2% off).
  • The Verdict: As expected, the "Connect-the-Dots" guesses (Stage 1) were more accurate than the "Growth Chart" guesses (Stage 2). This proves that the two-stage order (try to connect dots first, use charts only if necessary) is the right approach.

Important Limitations (What the Paper Says)

The paper is very honest about what it doesn't do yet:

  • It's a "Single Guess": The method gives one best guess for a missing number. It doesn't calculate the "uncertainty" or the range of possibilities. If you use these guesses in a final medical study, you might accidentally think you know more than you actually do.
  • The "Stability" Assumption: Stage 2 assumes a child stays on the same growth track. If a child has a sudden illness or a growth spurt that changes their percentile drastically, this method might not catch that shift. However, because of the "labels," researchers can spot these guesses and treat them with caution.
  • Synthetic Only: These results are based on computer simulations, not real patients. The method works perfectly in the simulation, but it needs to be tested on real-world messy data before being used for serious medical conclusions.

Summary

This paper offers a simple, transparent tool to fix missing growth data. It prioritizes using a person's own history first, then falls back to standard growth charts if needed, while loudly announcing which "standard" was used for every single calculation. It's a way to make sure that when we combine data from different sources, we aren't accidentally mixing apples and oranges without knowing it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →