← Latest papers
📊 statistics

Considerations for the Integration of Randomized Controlled Trials and Real-World Data

This article outlines a principled framework for integrating randomized controlled trials and real-world data to support individualized clinical decision-making, detailing how distinct objectives shape design and analytic choices while addressing practical challenges like data curation, comparability, and sensitivity analysis to maintain regulatory-grade evidentiary standards.

Original authors: Sky Qiu, Charles Barr, Lauren Dang, Issa Dahabreh, Larry Han, Kajsa Kvist, Hana Lee, Andrew Mertens, Nerissa Nance, Lei Nie, Kara Rudolph, Xu Shi, Jens Tarp, Salina P. Waddy, Kenneth Wiley, Andy Wilso
Published 2026-04-14
📖 6 min read🧠 Deep dive

Original authors: Sky Qiu, Charles Barr, Lauren Dang, Issa Dahabreh, Larry Han, Kajsa Kvist, Hana Lee, Andrew Mertens, Nerissa Nance, Lei Nie, Kara Rudolph, Xu Shi, Jens Tarp, Salina P. Waddy, Kenneth Wiley, Andy Wilson, Margot Lisa Jing Yann, Zhiwei Zhang, Tianyue Zhou, Maya Petersen, Mark van der Laan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a chef trying to perfect a new recipe for a soup that will feed a whole city.

The Gold Standard: The Controlled Kitchen (RCT)
Traditionally, to prove your soup is delicious and safe, you run a Randomized Controlled Trial (RCT). This is like a high-end, sterile test kitchen. You pick a very specific group of people (say, only left-handed chefs aged 30-40), give them strict instructions, control the temperature of the stove, and measure every ingredient precisely.

  • The Good: Because you control everything, you know exactly that the soup tastes good. The result is scientifically "true" for that specific group.
  • The Bad: The test kitchen is small, expensive, and the people in it aren't like the average person in the city. Maybe the city has many elderly people, or people who cook in small apartments with weak stoves. Your "perfect" soup might taste different (or even be unsafe) when 10,000 regular people try to make it in their own messy kitchens.

The Real World: The Street Food Festival (RWD)
Then there is Real-World Data (RWD). This is like watching thousands of people make your soup at a chaotic street food festival. They use different pots, forget ingredients, turn the heat up and down, and have all sorts of dietary restrictions.

  • The Good: It shows you how the soup actually performs in the real world. It covers a huge variety of people and long-term effects.
  • The Bad: It's messy. You don't know if they liked the soup because of your recipe, or because they added extra salt, or because they were just hungry. It's hard to tell cause from effect.

The Big Idea: Mixing the Two
This paper is a guidebook for mixing the Test Kitchen with the Street Festival. The goal is to get the scientific rigor of the test kitchen plus the real-world relevance of the festival.

The authors (a team of statisticians, doctors, and regulators) say: "Don't just look at the test kitchen results and guess how they apply to the world. Don't just look at the festival and guess if the recipe is good. Integrate them."

Here is how they suggest doing it, using simple analogies:

1. Why Mix Them? (The Objectives)

The paper lists five reasons to mix the data, like adding different ingredients to a stew:

  • More Power (Bigger Sample): If your test kitchen only had 10 people, you might miss a rare side effect (like a stomach ache that happens to 1 in 1,000). The street festival has 10,000 people. Combining them helps you spot those rare events.
  • Generalizability (The "Average" Person): Your test kitchen only had 30-year-olds. The street festival has kids, seniors, and people with different health issues. By mixing them, you can say, "This soup works for everyone, not just 30-year-olds."
  • Transporting (The "Different" City): Maybe your test kitchen was in New York, but you want to know if the soup works in Tokyo (different climate, different ingredients). You use the festival data from Tokyo to "transport" your findings there.
  • Long-Term Watch (The "Time" Machine): Test kitchens usually only run for 6 months. The street festival has been going on for 10 years. You can link your test kitchen participants to the festival records to see what happens to them 5 or 10 years later.
  • When You Can't Have a Control Group: Sometimes, for rare diseases, you can't find enough people for a test kitchen. You might only have one group of patients. In this case, you use the "festival" data as a fake control group to compare against.

2. The Danger Zone: The "Apples and Oranges" Problem

The paper warns that you can't just dump the festival data into the test kitchen data. That's like trying to compare a Michelin-starred chef's knife skills with a toddler's butter knife. They aren't the same!

The "Misalignment" Checklist:
Before mixing, you must check for mismatches:

  • Time Mismatch: Did the festival happen 20 years ago when medicine was different? (Like comparing soup made with 1990s ingredients to 2024 ingredients).
  • Definition Mismatch: In the test kitchen, "eating the soup" means finishing the bowl. In the festival, does "eating" mean taking one spoonful? You need to make sure you are measuring the same thing.
  • Selection Bias: In the festival, maybe only the people who liked the soup stayed to fill out a survey. If you only look at them, you think the soup is perfect. You have to be careful not to pick only the "happy" customers.

3. The Recipe for Success (The Causal Roadmap)

The authors propose a step-by-step "Causal Roadmap" (a fancy term for a recipe) to ensure the mix works:

  1. Ask the Right Question: Are you asking "Does it work in the test kitchen?" or "Does it work for the whole city?"
  2. Define the Target: Be specific about who you are talking about.
  3. Check the Assumptions: Be honest about what you are guessing. "We assume the festival people are similar enough to the test kitchen people if we account for their age and health."
  4. Run Sensitivity Checks: This is like tasting the soup at different stages. "What if my assumption about age was wrong? Does the result change?" If the result flips completely, your recipe is shaky.

4. The "Magic" Tools (Estimators)

The paper mentions fancy statistical tools (like "Bayesian borrowing" or "Conformal prediction"). Think of these as smart filters.

  • Instead of blindly mixing all the festival data, these tools act like a sieve. They say, "We will only borrow data from the festival customers who look very similar to our test kitchen customers."
  • This prevents the "noisy" festival data from ruining the "clean" test kitchen data.

The Bottom Line

This paper is a plea for honesty and precision.
It tells regulators and scientists: "We have amazing tools to mix the perfect test kitchen with the messy real world. But we must do it carefully. We can't just throw data together and hope for the best. We need to define our goals, check our ingredients, and admit where we are making guesses."

If done right, this integration gives us better medicine: treatments that are proven to work in the lab and proven to work for real people in real life.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →