Using large language models for sensitivity analysis in causal inference: cases studies on Cornfield inequality and E-value
This study demonstrates that large language models, when guided by structured prompts, can accurately calculate E-values, interpret robustness to unmeasured confounding, and suggest plausible confounders, thereby serving as effective tools to assist clinicians and researchers in conducting sensitivity analyses for observational studies.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery: Did smoking cause lung cancer, or was it just a coincidence caused by something else we didn't notice?
In the world of science, this is called an observational study. Researchers watch people in the real world (they can't force people to smoke or not smoke like in a lab experiment) and look for patterns. But here's the catch: there might be a "ghost" in the room—an unmeasured confounder. This is a hidden factor (like genetics, pollution, or stress) that affects both the cause (smoking) and the effect (cancer), making it look like smoking is the culprit when it might not be.
To catch this "ghost," scientists use a special tool called Sensitivity Analysis. Think of it as a stress test for their conclusions. They ask: "How strong would this hidden ghost have to be to completely ruin our conclusion?"
Two famous tools for this stress test are the Cornfield Inequality and the E-value.
- The E-value is like a "Monster Strength Meter." If the E-value is high (say, 10), it means the hidden ghost would need to be a 10-foot-tall, fire-breathing dragon to explain away the results. Since 10-foot dragons don't exist, the scientists can be confident smoking really does cause cancer. If the E-value is low (say, 1.2), the ghost only needs to be a small, invisible mouse to ruin the conclusion, so the scientists should be very suspicious.
The Problem
Calculating these numbers and interpreting what they mean is like trying to solve a complex algebra equation while juggling. It's hard for doctors and researchers who aren't math wizards. They need a helper.
The New Helper: The AI Detective
This paper asks a big question: Can Artificial Intelligence (specifically Large Language Models like ChatGPT, Claude, and Gemini) act as that helper?
The researchers treated four different AI models like new interns. They gave them four real-life mystery cases (smoking, back pain, Alzheimer's, and environmental pollution) and said:
- Do the Math: Calculate the "Monster Strength Meter" (E-value) for us.
- Give an Opinion: Based on that number, is the result safe from hidden ghosts, or is it shaky?
- Spot the Ghost: Can you guess what hidden factors we might have missed?
The Results: Who Passed the Test?
The researchers put the AI interns through the wringer, and here is what happened:
1. The Math Whizzes (ChatGPT, Claude, Gemini)
These three AIs were like super-accurate calculators. They looked at the data and calculated the "Monster Strength Meter" perfectly, matching the numbers the original human scientists had found. They didn't need the formula written out; they just knew how to do it.
2. The "One-Size-Fits-All" Intern (DeepSeek)
This model was a bit off. It tried to do the math but made small mistakes, like a calculator with a sticky button. It wasn't terrible, but it wasn't perfect.
3. The Storytellers (Interpretation)
When it came to explaining what the numbers meant, the top three AIs were great.
- If the "Monster Strength" needed was huge, they said, "Don't worry, this result is solid."
- If the strength needed was tiny, they said, "Be careful, a small ghost could ruin this."
- The nuance: Interestingly, the original human papers often gave a single, vague answer for all their data. The AIs, however, were more flexible. They said, "Well, for this specific case, it's solid, but for that other case, it's shaky." They were actually more careful than the humans!
4. The Brainstormers (Suggesting Hidden Ghosts)
The AIs were asked to guess what hidden factors might be lurking. They did surprisingly well.
- For a smoking study, they suggested things like "genetics" or "dust in the workplace."
- For a back pain study, they suggested "sleep quality" or "job stress."
- These weren't random guesses; they were logical, scientifically sound ideas that real researchers might have missed.
The Big Takeaway
This paper is like a report card for AI in the scientific world. It shows that if you give an AI a clear set of instructions (a "structured prompt"), it can:
- Do the heavy math lifting.
- Explain the results in plain English.
- Help researchers spot the "ghosts" (hidden factors) they might have missed.
In short: AI isn't replacing the detective yet, but it's becoming an excellent sidekick. It helps scientists double-check their work, ensuring that when they say "A causes B," they aren't just being fooled by a hidden variable. It makes the whole process of scientific discovery a little less scary and a lot more reliable.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.