← Latest papers
📊 statistics

Sample size and power determination for assessing overall SNP effects in joint modeling of longitudinal and time-to-event data

This paper addresses the gap in genetic study design by deriving a closed-form sample size formula for assessing overall SNP effects within a joint modeling framework of longitudinal biomarkers and time-to-event outcomes, validating its accuracy through simulations and a real-world application.

Original authors: Yuan Bian, Shelley B. Bull

Published 2026-02-18
📖 5 min read🧠 Deep dive

Original authors: Yuan Bian, Shelley B. Bull

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery: Why do some people with Type 1 diabetes develop serious eye or kidney problems, while others don't?

You suspect two things are happening:

  1. The Direct Clue: A specific genetic "glitch" (a SNP) in their DNA might directly cause the damage.
  2. The Indirect Clue: That same genetic glitch might make their blood sugar levels (HbA1c) fluctuate wildly over time, and those wild swings are what actually cause the damage.

To solve this, you need to look at two types of evidence simultaneously:

  • The Longitudinal Evidence: Repeated blood sugar tests taken over years (like a diary of their health).
  • The Time-to-Event Evidence: The exact moment a complication (like retinopathy) appears.

The Problem: "How Many Detectives Do We Need?"

In scientific studies, you can't just guess how many people to recruit. If you pick too few, you might miss the answer even if it's right there (a false negative). If you pick too many, you waste money and time.

The authors of this paper, Yuan Bian and Shelley Bull, realized that while we have great tools to analyze this data, we didn't have a good "recipe" to calculate how many people we need before we start the study. Existing recipes were designed for simple drug trials (like comparing a pill to a placebo), but genetics is trickier because:

  • Genes aren't just "on" or "off"; they come in three versions (0, 1, or 2 copies).
  • Genes can be helpful or harmful (unlike a drug which is usually just one or the other).

The Solution: A New "Magic Formula"

The authors created a closed-form sample size formula. Think of this as a specialized calculator for genetic detectives.

Here is how the formula works, using a simple analogy:

Imagine you are trying to hear a whisper (the genetic effect) in a noisy room (biological variation).

  • The Volume of the Whisper: This is the Effect Size. If the gene has a huge impact, you need fewer people to hear it. If the impact is tiny, you need a huge crowd.
  • The Clarity of the Room: This is the Allele Frequency. If the gene is very rare (like a rare accent), it's harder to hear, so you need more people. If it's common, it's easier.
  • The Duration of the Listening: This is the Follow-up Time. The longer you listen (follow the patients), the more likely you are to catch the whisper.

The authors' formula combines the direct whisper (the gene acting alone) and the indirect whisper (the gene acting through blood sugar levels) to tell you exactly how many "ears" (participants) and how many "events" (complications) you need to be 90% sure you'll hear the truth.

The "Two-Stage" Shortcut

Calculating this is usually a nightmare for computers because the math is incredibly complex (like trying to solve a Rubik's cube while juggling). The authors suggest a clever shortcut called a Two-Stage Method:

  1. Stage 1: First, look at the blood sugar data alone to see how the gene affects it.
  2. Stage 2: Remove that genetic influence from the blood sugar data, then look at the remaining data to see how it affects the complications.

This is like peeling an onion: you remove the outer layer (the gene's effect on blood sugar) to see the core (the gene's total effect on the disease) without getting overwhelmed by the math.

The Real-World Test: The DCCT Study

To prove their formula works, they applied it to a famous real-world study called the Diabetes Control and Complications Trial (DCCT).

  • The Setup: This study had about 1,400 people split into two groups: one getting "intensive" care (very strict blood sugar control) and one getting "conventional" care.
  • The Result: When they ran their new formula on this data, the result was a bit of a shocker. The study wasn't big enough!
  • The Lesson: Even though the DCCT was a massive, famous study, it didn't have enough people to detect small genetic effects with the high standards required for modern genetics (where you need to be extremely sure you aren't just seeing random noise).

Why This Matters

This paper gives researchers a blueprint. Before they spend millions of dollars and years of time collecting data, they can use this formula to ask:

  • "If I want to find a gene that has a small effect, do I need 500 people or 5,000?"
  • "If I only follow patients for 5 years instead of 10, will I lose my ability to find the answer?"

In short: The authors built a better ruler for measuring genetic mysteries. They showed us that without this ruler, we might be building houses on shaky ground, and they proved that even famous studies might need to be bigger than we thought to solve the genetic puzzle of diabetes.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →