← Latest papers
📄 medicine

How Much Non-response Matters? Diagnosis Prevalence Estimates in Survey Participants Versus the Total Sample

This study of German health insurance data demonstrates that while non-response bias initially distorts disease prevalence estimates, statistical weighting for demographics generally minimizes these differences for most conditions, though significant biases persist for specific mental disorders and dementia.

Original authors: Timm Frerk, Felicitas Vogelgesang, Roma Thamm, Thom Julia, Saam Joachim, Catharina Schumacher, Ursula Marschall, Thomas G. Grobe

Published 2026-08-25
📖 1 min read☕ Coffee break read

Original authors: Timm Frerk, Felicitas Vogelgesang, Roma Thamm, Thom Julia, Saam Joachim, Catharina Schumacher, Ursula Marschall, Thomas G. Grobe

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Technical Summary: Diagnosis Prevalence Estimates in Survey Participants Versus the Total Sample

Problem Statement
Population-based surveys are a primary source for estimating disease prevalence, yet declining response rates over recent decades have raised significant concerns regarding non-response bias. This bias occurs when the probability of participation is associated with health status or diagnosis frequencies. While previous studies have compared health claims data and survey data, they often relied on unrelated samples or were restricted to respondents only, limiting their ability to quantify how study participation itself distorts prevalence estimates. There is a need for a direct assessment of non-response bias using diagnostic information for the entire invited sample, rather than just the participating subsample, to determine the generalizability of survey-based estimates, particularly in settings with modest response rates.

Methodology
The study utilized data from the "OptDatPMH" project, linking survey data with health claims data from BARMER, a large German statutory health insurance fund.

  • Study Population: A stratified random sample of 26,000 individuals (aged ≥18) was drawn from the BARMER insured population (nationally representative across 512 strata defined by sex, 5-year age groups, and federal state).
  • Survey Procedure: Participants were invited to complete a written health questionnaire (18 pages, ~20–30 minutes) in October 2021. After two reminders, 7,110 questionnaires were returned (27.3% response rate). After exclusions for invalid responses, the final analytic total sample (TS) comprised 25,977 individuals, with 7,087 forming the participating subsample (PS).
  • Data Sources:
    • Survey Data: Self-reported diagnoses and sociodemographics.
    • Health Claims Data: ICD-10 codes from outpatient and inpatient records for the four quarters preceding the survey (Q4/2020–Q3/2021). A diagnosis was considered present if documented as a confirmed diagnosis in at least one quarter (M1Q criterion).
  • Statistical Analysis: Prevalence estimates were calculated for both the TS and PS. The study employed population-based weights reflecting the 2021 German population distribution (sex, age, region) to adjust both samples, as the exact final population distribution was unknown at the time of sampling.
    • Differences were tested using chi-square tests for weighted data.
    • To visualize deviations without overemphasizing rare diseases, the authors introduced the "Difference as a Percentage of the Total Mean" (DPTM) metric, expressed as a percentage of the overall prevalence in the total sample.
    • Analyses were conducted at the level of ICD chapters, ICD groups, and nine selected specific diagnoses.

Key Contributions

  • Direct Comparison Framework: The study provides a rare, direct comparison of diagnosis prevalence between survey participants and the total invited sample using a common sampling frame and linked claims data, overcoming the limitations of previous studies that used unrelated samples.
  • Quantification of Bias: It quantifies the extent to which selective participation distorts prevalence estimates across a wide range of diagnostic categories (ICD chapters and groups) after standard demographic weighting.
  • Metric Development: The introduction of the DPTM metric offers a balanced approach to visualizing absolute and relative deviations across subgroups, avoiding the pitfalls of focusing solely on relative differences in low-prevalence subgroups.

Results

  • Participation Patterns: Participation was higher among women (30.1%) than men (24.5%) and increased with age (37.9% for ≥65 years vs. 18.2% for <45 years).
  • Aggregate Level (ICD Chapters):
    • Unweighted estimates showed significant differences between participants and the total sample for most ICD chapters, with participants generally showing higher prevalence.
    • After weighting for sex, age, and region, differences were generally small at the aggregate level. However, 11 of 22 ICD chapters still showed statistically significant differences. In 10 of these, prevalence remained higher among participants.
    • Exception (Mental Disorders): ICD Chapter V (Mental and behavioural disorders) was the primary exception, showing lower prevalence estimates among participants compared to the total sample.
  • Specific Diagnoses:
    • Mental Health: Significant underestimation was found for participants regarding organic mental disorders (F00-F09), substance use disorders (F10-F19), and mood disorders (F30-F39).
    • Selected Conditions: Of nine selected diagnoses, three showed significant differences: dementia (participants: 0.6% vs. total: 1.4%), alcohol-related disorders (1.5% vs. 1.9%), and anxiety disorders (6.8% vs. 7.6%). Dementia showed the most pronounced deviation.
    • Somatic Diseases: After weighting, none of the selected somatic diseases (e.g., diabetes, obesity, hypertension, heart failure) showed statistically significant differences between the total sample and participants.
  • Subgroup Analysis: Heatmaps of DPTM values confirmed a general tendency toward overestimation in somatic chapters across most age/sex subgroups, with notable exceptions for mental disorders and specific subgroups (e.g., women ≥70 in blood disorders, men in congenital malformations).

Significance and Claims
The paper concludes that in this specific setting, selective participation had only a limited effect on prevalence estimates for many common diagnoses after standard demographic weighting (sex, age, region).

  • Robustness of Surveys: Despite a modest response rate of 27.3%, survey-based prevalence estimates for many common somatic diagnoses appear more robust than often assumed. The findings suggest that declining response rates do not inevitably compromise epidemiological validity for these conditions.
  • Role of Weighting: Standard weighting accounts for a substantial share of participation-related distortion, though diagnosis-specific bias remains.
  • Vulnerable Groups: The study highlights that diagnosis-specific non-response bias remains a relevant concern for certain mental disorders (particularly dementia and substance use) and vulnerable subgroups. For these conditions, lower participation among affected individuals leads to underestimation that weighting does not fully correct.
  • Implication: The results challenge the general conclusion that low response rates automatically invalidate population-based surveys. However, they caution that for specific mental health conditions and vulnerable populations, bias remains a critical issue that requires careful interpretation.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →