← Latest papers
📄 infectious diseases

Simulation of synthetic health records for assessment of causal inference methods for vaccine efficacy

This paper demonstrates that low-fidelity synthetic datasets, generated from the EAVE-II platform using marginal structural models to simulate various confounding scenarios, can effectively validate causal inference methods for vaccine efficacy assessment, particularly showing that inverse probability of treatment weighting is necessary to recover true causal parameters under strong confounding.

Original authors: Velasco Pardo, V., Daines, L., Katikireddi, S. V., Ritchie, L., Robertson, C., Simpson, C. R., McCowan, C., Swallow, B.

Published 2026-07-20
📖 4 min read☕ Coffee break read

Original authors: Velasco Pardo, V., Daines, L., Katikireddi, S. V., Ritchie, L., Robertson, C., Simpson, C. R., McCowan, C., Swallow, B.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are a detective trying to solve a mystery: Did a new vaccine actually stop people from getting sick, or did the sick people just happen to be the ones who were already older or sicker to begin with? In the world of science, this is called "causal inference." It's the art of figuring out if A really caused B, rather than just happening at the same time. Usually, the best way to solve this is a "randomized controlled trial," where you flip a coin to decide who gets the medicine and who doesn't. But in real life, especially during a fast-moving pandemic, you can't always flip coins; you have to look at the messy, real-world records of millions of people. The problem is, those records are full of "confounders"—clues that mix things up, like age or health history, making it hard to see the true picture. To test if their detective tools work, scientists usually need to peek at the real case files, but privacy laws lock those files away tight. So, they need a way to practice their detective skills without breaking the rules.

This is where the story gets clever. A team of researchers decided to build a "fake universe" of health records. Think of it like a flight simulator for doctors and data scientists. Instead of flying a real plane through a storm (which could be dangerous and requires a real pilot's license), they built a computer program that creates a perfect storm with known wind speeds and turbulence. They can crash the plane a hundred times in the simulator to see if their navigation tools work, all without anyone getting hurt or needing a real airport. In this specific study, the researchers created a digital twin of the Scottish population's health data. They didn't use real people's secrets; instead, they used the "shape" of the real data (how many men, how many women, how many people in each age group) to build a fake population of 100,000 digital citizens. They programmed this fake world with a "ground truth"—they secretly decided exactly how effective the vaccine was and exactly how much the confounders messed things up. Then, they let their statistical tools try to find that truth.

The researchers tested two different detective tools on their fake data. The first tool was a simple approach that just looked at the numbers without adjusting for the messy background clues. The second tool was a more complex method called "Inverse Probability of Treatment Weighting" (IPTW), which is like giving extra weight to the clues from people who were unusual (like a young person getting a vaccine when most young people didn't) to balance out the picture. When the "messiness" in their fake world was low, both tools did a great job finding the right answer. But as they cranked up the confusion—making the confounding factors stronger and stronger—the simple tool started to get the answer wrong, missing the true effect of the vaccine. The fancy weighted tool, however, kept its cool. Even when the fake world was chaotic and confusing, the weighted tool managed to recover the true "ground truth" they had programmed in.

The paper shows that while simple math can work when things are easy, it starts to fail when the real world gets complicated. The weighted method proved to be much more reliable in these simulations, successfully cutting through the noise to find the real cause-and-effect relationship. However, the authors are very clear that this is a test drive, not a final destination. They simulated 100,000 people across five different types of rare health outcomes and five different levels of "confusion," but they only simulated one visit to the clinic, not a long series of check-ups over years. They also emphasize that these synthetic records are not real; they are a training ground. You can't use them to make medical decisions for real people because they don't contain the actual, messy details of individual lives. Instead, they are a powerful new playground. They allow researchers to build and test their analysis pipelines before they are granted access to the real, sensitive data. This means that when the real data finally arrives, the scientists won't be starting from scratch; they'll already have their tools polished and ready to go, saving precious time in the race to understand how vaccines work in the real world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →