Auditing Institutional Heterogeneity for Generative AI in Patient Education: A Large-Scale Study of 102 US Transplant Handbooks
This large-scale study audits 102 US transplant center handbooks and finds that institutional editorial voice creates greater agreement than organ-specific consistency, while revealing critical information gaps—particularly regarding reproductive health—that pose significant risks for the deployment of grounded generative AI in patient education.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are walking into a massive library where every book is supposed to teach you how to care for a very special, fragile garden. But here's the twist: this isn't just one library; it's a collection of 102 different guidebooks from 23 different gardeners across the country. In the world of medicine, these "gardens" are human organs, and the "gardeners" are transplant centers. For years, doctors have been building smart, AI-powered assistants to answer patients' questions using these guidebooks. The big idea was simple: if the AI reads the local guidebook, it will give advice that matches exactly what that specific hospital does. It's like having a tour guide who knows the secret paths of your specific park.
But there was a hidden worry that no one had checked on a large scale: Do all these guidebooks actually agree with each other? Imagine if one book says "water the flowers every morning" and another says "only water them on Tuesdays." If the AI picks the wrong book, the garden could suffer. This paper dives into that exact question. It treats the collection of medical handbooks like a giant puzzle, checking millions of pairs of pages to see if the advice is consistent or if the "voice" of the hospital matters more than the type of organ being discussed. The goal is to make sure that when a patient asks a smart computer for help, they get safe, reliable answers that don't accidentally contradict their doctor's plans.
The Great Handbook Detective Story
So, what did the researchers actually do? They acted like super-powered detectives, but instead of solving crimes, they solved a mystery of consistency. They gathered 102 patient-education handbooks from 23 US solid-organ transplant centers. These centers deal with hearts, lungs, kidneys, livers, and pancreases. They paired these handbooks with 1,115 real questions that patients actually ask, like "Can I travel?" or "When can I have a baby?"
Then, they unleashed a "judge" (a very advanced AI) to read every possible combination of answers. They didn't just look at a few examples; they ran a massive audit of 5,730,465 pairwise comparisons. That's over 5.7 million checks to see if Handbook A and Handbook B were saying the same thing, different things, or nothing at all.
The Big Surprises
Here is what the detectives found, and it's a bit more complicated than "everyone agrees."
1. The "House Style" is Stronger Than the Organ
You might think that all heart doctors would agree with each other, and all kidney doctors would agree with each other, regardless of where they work. The study found that this isn't quite true. In fact, the "editorial voice" of a single hospital is so strong that two handbooks from the same hospital (even if one is for hearts and one is for lungs) agreed with each other more than two handbooks for the same organ from different hospitals. It's like if a bakery in Chicago always puts extra sprinkles on every cake, whether it's a chocolate or a vanilla one, while a bakery in New York never uses sprinkles at all. The "Chicago style" matters more than the "chocolate vs. vanilla" difference. This means that if you mix handbooks from different hospitals into one AI, you might get confused advice just because of the hospital's unique writing style.
2. The "Silent" Topics Hit the Hardest
The biggest problem wasn't that the handbooks disagreed; it was that they often said nothing. The study found that 78.9% of the time, at least one handbook in a pair simply didn't have an answer. But here is the sad part: the topics that were missing the most were the ones that mattered most to people who are often overlooked.
- Reproductive health (questions about pregnancy and having babies) was the single most silent topic, with 82% of handbooks not addressing it at all.
- Yet, when two handbooks did answer it, they disagreed on high-stakes details 86% of the time.
- This is a "double jeopardy": patients are most likely to get no answer on this topic, but if they do get an answer from two different sources, they are most likely to get conflicting advice.
- Other missing topics included mental health, financial struggles, and care for special populations like children or the elderly.
3. The "High-Stakes" Arguments
When the handbooks did disagree, it wasn't usually about small things like "what to eat for breakfast." The disagreements clustered into 991 different themes, but the most dangerous ones were about pregnancy timing and immunosuppression (the drugs patients take to stop their body from rejecting the new organ). For example, one handbook might say you can try for a baby six months after surgery, while another says wait a year. These aren't just typos; they are life-altering decisions.
4. We Can Predict the Trouble
The researchers also built a crystal ball of sorts. They found that they could predict which questions would cause the most disagreement just by looking at how the question was asked. Questions starting with "Can I..." or questions about travel were the most likely to trigger conflicting answers. The model was good at this, with a score (AUC) of 0.77, meaning it could spot the troublemakers before the AI even tried to answer them.
Why This Matters for You
The paper concludes that we can't just plug any pile of hospital handbooks into a smart AI and expect it to work perfectly. The "house style" of the hospital creates its own reality, and the most vulnerable patients are the ones getting the least information.
The researchers suggest a few fixes:
- Audit before you launch: Hospitals should check their own handbooks for gaps and contradictions before letting an AI talk to patients.
- The "Can I?" Filter: If a patient asks a "Can I travel?" or "Can I do X?" question, the system should be extra careful, maybe even showing a human doctor first, because those are the questions where the handbooks disagree the most.
- Honesty about gaps: If the AI doesn't know the answer (especially about pregnancy or mental health), it should say "I don't know" or "This varies by hospital" rather than making something up or giving a generic answer that might be wrong for that specific patient.
In short, this study shows that while AI is a powerful tool, it needs to be taught that not all hospitals speak the same language, and that some of the most important questions are the ones that are currently being left unanswered.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.