← Latest papers
🤖 AI

Protecting patient privacy in clinical foundation models: Technical and legal perspectives

This paper proposes a practical, context-aware framework to assess and mitigate privacy risks in clinical foundation models by addressing model-mediated data leakage through complementary technical and legal strategies, as existing regulations like HIPAA and GDPR offer limited guidance for these indirect threats.

Original authors: Sana Tonekaboni, Lena Stempfle, Sasha Ronaghi, Corinna Coupette, I. Glenn Cohen, Emily Alsentzer, Marzyeh Ghassemi

Published 2026-08-11
📖 6 min read🧠 Deep dive

Original authors: Sana Tonekaboni, Lena Stempfle, Sasha Ronaghi, Corinna Coupette, I. Glenn Cohen, Emily Alsentzer, Marzyeh Ghassemi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you've built a super-smart robot chef. You fed it millions of recipes and cooking videos from around the world so it could learn to make anything from soup to soufflé. Now, you want to let people use this robot to help them cook. But here's the catch: because the robot learned so deeply, it might accidentally remember and serve up a specific, secret family recipe from a stranger it was trained on, or even reveal that a specific person's favorite dish was in its training book. This isn't just about the robot stealing a file; it's about the robot acting in a way that leaks secrets it wasn't supposed to know. This is the world of "foundation models" in healthcare. These are giant AI systems trained on massive amounts of patient data—like medical notes, X-rays, and lab results—to help doctors diagnose diseases, predict health risks, and plan treatments. While these tools promise to revolutionize medicine, they come with a sneaky new kind of privacy risk. It's not just about hackers breaking into a database; it's about the AI itself "leaking" sensitive details through its answers, even if the original data was scrubbed of names and addresses. The big question is: how do we keep these powerful tools safe without locking them away?

This paper, written by a team of researchers from places like MIT, Stanford, and Harvard, tackles exactly that problem. They argue that current privacy rules, like the ones in the US (HIPAA) and Europe (GDPR), are like old traffic laws written for cars, but we are now driving self-driving rockets. These laws focus on protecting the data before it goes into the AI, but they don't really know how to handle the weird ways an AI might accidentally spill the beans after it's been trained.

To fix this, the authors propose a new way of thinking about privacy risk using a simple two-part map. Imagine a graph where one side asks, "How much does the person asking the question already know?" and the other side asks, "How much sensitive info does the AI give away?"

  • The "Prior Knowledge" Axis: Sometimes, you don't need to know anything to trick the AI into spilling a secret (like asking it to "make up a story" and it accidentally tells a real patient's story). Other times, you need to give the AI a tiny clue, like a specific medication name or a rare symptom, to get it to reveal more details about a specific person.
  • The "Leaked Info" Axis: On the other side, the AI might just give you a piece of information that could be linked to someone (like a rare disease pattern), or it might give you information that directly identifies them (like a name or a specific phone number).

The authors use this map to look at six different "leakage scenarios" to show how things can go wrong. For example:

  1. The "Copycat" Scenario: An AI trained on brain scans is asked to generate a generic image, but it accidentally spits out a near-perfect copy of a real patient's scan, complete with unique quirks that could identify them.
  2. The "Memory Lane" Scenario: A user asks an AI to write a hospital note for a specific type of patient, and the AI, having memorized a real note from its training, writes almost the exact same story, revealing private details the patient never shared publicly.
  3. The "Autocomplete Trap": A doctor is typing a note for Patient A, and the AI suggests a follow-up plan that includes the name and phone number of Patient B because it got confused and pulled from the wrong memory.
  4. The "Membership Leak": A researcher asks the AI a question and notices its behavior is weirdly consistent with a specific group of patients, proving that those specific people's data was used to train the model, even if no names are revealed.

The paper suggests that these risks are real and happen in both "local" settings (where a hospital keeps the AI on its own secure servers) and "public" settings (where anyone can use the AI online). They find that current laws are a bit fuzzy here. For instance, in the US, if a hospital removes names from data before training an AI, they might think they are safe under HIPAA. But the authors point out that if the AI can still be tricked into revealing that specific person's data, the hospital might actually have "actual knowledge" that the data isn't truly anonymous, which could be a legal problem they didn't expect. Similarly, in Europe, the rules say data must be "identifiable" to be protected, but the authors argue that even if you can't name the person immediately, if the AI reveals a rare combination of symptoms that could lead to identification, it's still a privacy breach.

So, what's the solution? The authors don't claim to have a magic wand that fixes everything instantly. Instead, they suggest a mix of technical and legal fixes. On the tech side, they recommend better testing (like "red teaming," where hackers try to break the AI to find leaks), using math to limit how much the AI can memorize (called "differential privacy"), and constantly checking the AI after it's released to make sure it hasn't started leaking secrets. On the legal side, they argue that we need to update our rules to look at the context of how the AI is used, not just the data itself. They suggest that companies and hospitals need to be more careful about who they share data with and how they train these models, perhaps requiring new kinds of contracts or oversight.

Ultimately, the paper suggests that we can't just rely on the old idea of "hiding names" to keep patients safe. As AI gets smarter and more integrated into our lives, we need a smarter, more flexible way to measure risk. It's like realizing that locking your front door is good, but you also need to check your windows, your basement, and make sure your neighbors aren't accidentally shouting your secrets. The authors hope that by using their new "risk map," we can keep the amazing benefits of medical AI while making sure patients' private lives stay private. They admit that the legal landscape is still shifting, with new laws like the EU's AI Act trying to catch up, but the core message is clear: we need to be proactive, not just reactive, when it comes to AI privacy.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →