← Latest papers
🧬 biology

IndepMR: An Ontology-Driven, Ancestry-Aware Interactive R Shiny Framework for Automated Two-Sample and Multivariable Mendelian Randomization with Mediation Analysis Using Public GWAS Summary Data

IndepMR is a novel, open-source R Shiny framework that automates two-sample, multivariable, and mediation Mendelian Randomization analyses by integrating ontology-driven trait retrieval and a rigorous ancestry-matching engine to ensure reproducible causal inference using public GWAS summary data.

Original authors: T. P.B. Yatawara¹, D. S. Wickramarachchi¹

Published 2026-08-18
📖 7 min read🧠 Deep dive

Original authors: T. P.B. Yatawara¹, D. S. Wickramarachchi¹

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

To understand how a specific habit or trait might cause a disease, scientists often face a tricky problem: people who have that trait might also share other hidden habits that actually cause the illness. Imagine trying to figure out if eating a certain food causes heart trouble, but the people who eat it also tend to smoke more. It becomes nearly impossible to tell which factor is the real culprit. For decades, researchers have used a clever workaround called Mendelian randomization to cut through this confusion. This method relies on the fact that our genes are shuffled randomly at conception, much like a lottery ticket, long before we are born and long before we develop any diseases. Because these genetic variations are assigned by chance, they act as natural, unbiased markers. If a person carries a genetic variant that naturally makes their body produce more of a certain substance, and that person is also more likely to develop a specific disease, scientists can infer that the substance itself is likely causing the disease, rather than some other lifestyle factor.

However, using this genetic lottery to solve real-world medical questions has become increasingly difficult as the amount of available data has exploded. Thousands of massive studies now exist, cataloging the genetic links to everything from height to diabetes, but they are scattered across different populations and described in confusing ways. A study might list its participants as "British," while another says "European," and a third might mix "Italian and European" in a single description. Without a way to automatically sort these labels, researchers might accidentally mix data from people who are genetically too similar, which skews the results, or they might miss crucial studies because they searched for the wrong word. A new tool called IndepMR, developed by researchers at the University of Colombo, aims to fix these bottlenecks. It is a computer program designed to act as a rigorous guide, automatically finding the right genetic studies, checking that the people in them are compatible, and running complex calculations to reveal true cause-and-effect relationships without requiring the user to be a coding expert.

The core innovation of this new tool lies in how it searches for information and how it checks the people behind the data. Traditionally, if a researcher wanted to study "high blood sugar," they might type those exact words into a database. But a study might have used the term "fasting glucose" or "HbA1c" instead, and a simple search would miss it entirely. IndepMR solves this by using a structured map of biological concepts, similar to a family tree for diseases and traits. If a user searches for "high blood sugar," the program understands that this is related to "fasting glucose" and "HbA1c" because they all branch from the same parent concept. This allows the tool to find every relevant study, even if the researchers used different names for the same condition. This approach proved powerful in testing the tool: when searching for the broad category of "lipids," the new tool found over 1,200 studies, whereas a standard keyword search found only 166.

Equally important is the tool's ability to check the ancestry of the people in the studies. For the genetic lottery to work fairly, the group of people providing the genetic clues for the exposure (like body weight) and the group providing the clues for the outcome (like diabetes) must be genetically compatible but not the exact same people. If they are too similar, the results can be biased; if they are from completely different ancestral backgrounds, the genetic markers might not work the same way. Existing tools often left this check to the human researcher, who had to manually read through text descriptions like "Italian, European" and guess if they matched. IndepMR automates this by breaking down these descriptions into a standardized hierarchy. It recognizes that "Italian" is a specific type of "European," and it flags when a study uses a broad label that might hide a specific overlap. This prevents researchers from accidentally comparing two groups that share the same underlying genetic pool, a mistake that could lead to false conclusions.

The researchers put their new system to the test by examining the link between body mass index and type 2 diabetes, a question that has been studied for years but yields different results depending on the population. They ran the analysis twice: once using data from people of European ancestry and once using data from people of Asian ancestry. In the European group, the tool confirmed what many expected: higher body mass index causes a significant increase in the risk of developing type 2 diabetes. The analysis was clean, with no signs of hidden errors or conflicting signals. However, the story changed when they looked at the Asian group. The tool initially found a result that seemed to suggest a protective effect, implying that higher body weight might lower diabetes risk. But the automated checks immediately raised a red flag. The data showed severe inconsistencies and signs that the genetic markers were picking up on other unrelated factors. The tool correctly identified that this surprising result was likely unreliable, driven by hidden biases in the data rather than a true biological reversal. This demonstrated that the tool does not just produce numbers; it acts as a quality control expert, telling researchers when a result looks suspicious.

To show the full power of the system, the team also demonstrated how it handles complex questions involving a middle step, or mediator. They tried to see if the effect of body weight on diabetes worked through a specific blood marker called HbA1c. The tool successfully guided them through the steps of combining three different sets of genetic data. However, it also caught a critical design flaw: the study for body weight and the study for the blood marker came from the exact same group of people. In a proper genetic study, these groups should be separate to ensure the results are not just a reflection of the same individuals' data. The tool flagged this overlap immediately, warning the researchers that any conclusion drawn from this specific combination would be biased. This feature is vital because it stops researchers from wasting time on flawed analyses or, worse, publishing results that look convincing but are actually artifacts of the study design.

The development of IndepMR represents a shift toward making sophisticated genetic analysis accessible to anyone with a scientific question, not just those who can write complex computer code. By automating the tedious and error-prone tasks of finding the right studies, checking their compatibility, and running the necessary statistical tests, the tool allows researchers to focus on the biology rather than the mechanics of the data. It enforces a level of rigor that is difficult to maintain manually, ensuring that the answers provided are based on solid, compatible evidence. While the tool cannot solve every problem—such as finding a perfect study when none exists—it provides a transparent, step-by-step environment where the limitations of the data are clearly visible. In the end, the goal is to make the process of discovering cause and effect in human health more reliable, ensuring that the conclusions drawn from our genetic lottery are as accurate as possible.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →