Eliciting associations between clinical variables from LLMs via comparison questions across populations
This paper proposes a method that uses structured patient comparison triplet questions and statistical modeling to elicit stable, clinically interpretable correlations from large language models, enabling invariant causal prediction to identify conservative candidate causal links across different patient subpopulations without accessing model internals.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a giant, super-smart librarian who has read almost every medical book, study, and paper ever written. You want to know how two specific health factors are connected—for example, "Does smoking more cigarettes usually mean your lungs are weaker?"
If you just ask the librarian directly, "How strong is the link between smoking and weak lungs?" they might give you a number. But there's a problem: the librarian might be guessing, remembering just one specific book, or even trying to please you by saying what they think you want to hear. They might also be bad at math, giving you a number that sounds right but isn't actually what they "know" deep down.
This paper proposes a clever, indirect way to ask the librarian questions to find out what they really know about these connections, without ever seeing their internal notes or brain.
The "Similarity Game" (Triplet Questions)
Instead of asking for a number, the researchers play a game of "Which one is more similar?"
Imagine you show the librarian three made-up patients:
- Patient A: Smokes 10 cigarettes a day, has average lung function.
- Patient B: Smokes 10 cigarettes a day, has very poor lung function.
- Patient C: Smokes 20 cigarettes a day, has unknown lung function.
You ask: "Is Patient C more similar to Patient A or Patient B?"
- If the librarian says "Patient A," they are ignoring the heavy smoking of Patient C.
- If the librarian says "Patient B," they are realizing that because Patient C smokes a lot, their lungs are likely worse, just like Patient B's.
By changing the numbers slightly (e.g., making Patient C smoke 15, then 25, then 30 cigarettes) and watching how the librarian's choice flips from A to B, the researchers can map out a "decision line." This line tells them how strongly the librarian connects smoking to lung health. It's like testing how much weight you need to put on a scale to tip the balance.
The "Different Neighborhoods" Trick (Prompted Environments)
The researchers realized that the librarian might have different "opinions" depending on the context. So, they created different "neighborhoods" or scenarios in their questions.
- Scenario 1: "Imagine a group of patients who are heavy smokers living in a city with smog."
- Scenario 2: "Imagine a group of patients who don't smoke but live in a very polluted area."
- Scenario 3: "Imagine a group of patients who don't smoke and live in clean, rural air."
They asked the similarity game in each of these three scenarios. If the librarian's "decision line" (how they connect the variables) stays the same no matter which neighborhood they are in, that suggests a strong, causal link. It's like saying, "No matter where you go, heavy smoking always leads to bad lungs."
If the line changes wildly between neighborhoods, it suggests the connection is weak or just a coincidence (a "spurious correlation").
What They Found
The researchers tested this on two real medical areas: COPD (a lung disease) and Multiple Sclerosis (a nerve disease).
- It Works Better Than Direct Questions: When they asked the librarian directly for numbers, the answers were all over the place—messy and inconsistent. But when they played the "Similarity Game," the answers were smooth, stable, and made medical sense.
- COPD Results: The method found strong, logical connections. For example, it confirmed that smoking history is linked to reduced lung diffusion (how well lungs swap oxygen), and that air pollution is linked to total lung volume.
- MS Results: The connections were weaker and harder to find, which the authors noted might be because the medical literature on MS is more scattered or complex than for COPD.
- Finding the "Real" Causes: By using the "Different Neighborhoods" trick, they could filter out the noise. They identified a small, very safe list of connections that seemed to be true causes (like smoking causing lung damage) rather than just random associations.
The Bottom Line
The paper doesn't claim to have discovered new medical facts. Instead, it shows a new tool for asking AI models questions.
Think of it like this: If you want to know what a person truly believes, don't just ask them "What do you think?" (they might lie or guess). Instead, ask them to make choices between options in different situations. By watching how they choose, you can figure out their true underlying beliefs.
The authors showed that by using this "choice-based" method, we can safely extract meaningful medical patterns from AI models without needing to see how the AI's brain works inside, and without getting tricked by the AI trying to please us.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.