Memorization in Large Language Models in Medicine: Prevalence, Characteristics, and Implications
This study systematically investigates the prevalence, characteristics, and implications of memorization in medical Large Language Models, revealing that it is significantly more common than in general domains, persists across different adaptation scenarios like continued pre-training and fine-tuning on real-world clinical data, and poses critical risks regarding privacy and generalizability.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a super-smart robot to become a doctor. You don't just give it a textbook; you feed it millions of pages of medical journals, patient notes, and exam questions. This is how "Large Language Models" (LLMs) learn medicine. Think of an LLM like a giant, eager student who reads everything in the library. Usually, we hope this student learns the concepts—understanding that a fever means infection, or that a specific drug treats a specific disease. But there's a tricky side to how these robots learn: sometimes, instead of truly understanding the lesson, they just memorize the exact words on the page, like a parrot repeating a sentence without knowing what it means. This is called "memorization." In the world of artificial intelligence, this is a double-edged sword. On one hand, remembering facts is good; on the other, if the robot memorizes a patient's private diary entry or a specific exam question's answer key, it could accidentally spill secrets or fail to think for itself when faced with a new situation. Scientists have been worried about this, but they didn't know exactly how bad it was for medical robots until now.
A team of researchers from Yale and other institutions decided to investigate this "parrot problem" in medical AI. They wanted to see how often these models just copy-paste what they've seen before, what kind of things they copy, and whether this copying is helpful or dangerous. They tested three different ways these models are trained: reading huge piles of medical books (continued pretraining), studying for specific exams (fine-tuning on benchmarks), and learning from real, messy hospital records (fine-tuning on clinical data).
The results were surprising and a bit alarming. The researchers found that medical AI models memorize way more than regular AI models. In fact, when these models were trained on medical data, they were regurgitating chunks of text they had seen before at rates between 10% and 20%, and sometimes even higher. To put that in perspective, if you asked a regular AI to finish a sentence, it might accidentally repeat a phrase from its training 1% of the time; the medical models did it 10 to 20 times more often.
The study broke this memorization down into three distinct flavors, like different types of snacks in a lunchbox:
- The Good Stuff (Beneficial): Sometimes the model remembers a clinical guideline or a drug interaction perfectly. This is helpful because it helps the AI give accurate advice.
- The Boring Stuff (Uninformative): Often, the model just repeats the same boring disclaimers or template phrases, like "This is an educational tool, not a medical device." This doesn't help it think; it just looks like it's talking.
- The Dangerous Stuff (Harmful): This is the scary part. The models were found to be memorizing and regurgitating sensitive patient information. In a test using over 13,000 real patient records from a hospital system, the AI reproduced 3,192 instances of private health information (PHI). This included names, dates, specific diagnoses, and even family relationships. Even worse, the researchers found that standard tools used to scrub private data from records often missed these secrets, meaning the AI was remembering things that humans thought were hidden.
One of the most interesting discoveries was how "sticky" this memory is. The researchers found that once a model memorized something during its initial "reading" phase, it didn't forget it when they later taught it new tasks. Up to 87% of the content memorized during the initial training remained even after the model was fine-tuned for new jobs. It's like if a student memorized a chapter of a book, and even after studying a whole new subject, they couldn't stop reciting that first chapter.
The team also noticed that memorization isn't just a result of training the model too long (overfitting). They watched the models learn over time and saw that the "copying" started happening very early in the training process, even while the model was still getting better at its actual job. This suggests that the way these models are currently built—simply feeding them more and more medical text without special safeguards—naturally leads to this high level of memorization.
The paper concludes that while these medical AI models are getting better at diagnosing diseases (their accuracy went up significantly after training), we can't ignore the privacy risks. The researchers suggest that we need to change how we build and test these models. We shouldn't just ask, "Is the diagnosis right?" We also need to ask, "Did it just copy a patient's secret?" They recommend using new techniques to stop the models from memorizing sensitive data and to encourage them to actually reason through problems rather than just repeating what they've read. The study doesn't say medical AI is broken, but it does sound a loud alarm that we need to be much more careful about what these digital doctors are remembering.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.