Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts
This study reveals that appending diversity, equity, and inclusion (DEI) prompts to medical queries causes large language models to frequently and inaccurately invent unrequested patient demographic attributes, a phenomenon termed "demographic injection" that significantly increases the rate of incorrect clinical recommendations across 47 models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the rapidly evolving field of artificial intelligence, a new generation of computer programs has learned to read and write like humans, offering advice on everything from literature to medicine. As these tools move into healthcare, experts have urged developers to program them with a specific kind of sensitivity: an awareness of diversity, equity, and inclusion. The goal is to ensure that when a doctor asks a computer for help with a patient, the machine considers how factors like race, income, or cultural background might influence health outcomes. This approach assumes that by explicitly asking the computer to think about these social realities, it will become fairer and more accurate. However, the relationship between a computer's instructions and its output is complex, and well-meaning prompts can sometimes trigger unexpected behaviors that distort the very information they are meant to clarify.
Researchers at Arizona State University investigated what happens when they apply these equity-focused instructions to medical questions. They tested forty-seven different artificial intelligence models, ranging from widely used commercial systems to specialized open-source programs, using a vast collection of medical multiple-choice questions. In their experiment, they took standard medical scenarios—descriptions of patients with specific symptoms but no mention of their background—and asked the models to solve them. They compared the models' answers under four different conditions: with no extra instruction, with a nonsensical instruction, with a neutral medical instruction, and with a single sentence directing the model to consider diversity and equity.
The results revealed a startling and consistent pattern. When the models received the equity-focused instruction, they frequently began to invent details about the patients that were never in the original description. The researchers call this phenomenon "demographic injection." In the control groups, where no special instruction was given, the models added unstated demographic details, such as a patient's race or social status, in less than one percent of cases. But when the equity prompt was added, that number jumped to thirty-three percent across every single model tested. This means that roughly one in three times the computer was asked to consider equity, it silently decided to assign a specific race, age, or economic situation to a patient who had none.
Most of the time, this invention was harmless. The models would add a general statement about how a certain disease affects a specific group of people in the real world, without changing their final medical recommendation. However, a smaller but significant portion of these responses went further. In about two and a half percent of the cases, the model attached the invented demographic directly to the specific patient in the story. This shift in the patient's identity caused the model to change its medical advice, and almost always, the new advice was wrong. For example, when presented with a case of a woman with a painless neck swelling and normal blood markers, a model might correctly identify the condition as a specific type of thyroid inflammation. But when prompted to consider equity, the same model might suddenly decide the patient was a Black or Hispanic woman, and based on that invented identity, switch its diagnosis to a different, incorrect disease.
The study found that the way the instruction was phrased mattered greatly. More detailed prompts that asked the model to consider social determinants of health or specific demographic groups caused the models to invent patient details even more often, with the rate of invention rising to as high as fifty-six percent. The researchers determined that this behavior was driven by the content of the equity instruction itself, not simply by the fact that the model was reading more text. They ruled out the idea that the models were merely distracted by extra words, as instructions that were just as long but unrelated to equity did not cause the same effect. Instead, the models appeared to have learned a strong association during their training: when they see words related to equity, they expect to see demographic details and feel compelled to provide them, even when the user has not asked for them.
This discovery highlights a subtle but critical flaw in how these systems process instructions. The computer is not acting out of malice or bias in the human sense; rather, it is following a pattern it has learned to be helpful. When told to consider equity, it interprets that as a command to fill in missing information, assuming that a complete medical picture requires a demographic label. In doing so, it rewrites the patient's identity, turning a generic case into a specific one that never existed. The researchers emphasize that while the overall accuracy of the models dropped only slightly, the errors were concentrated in the very cases where the models were trying to be most helpful. The study concludes that any instruction that nudges a model on how to reason can cause it to add unrequested details, and in the high-stakes world of medicine, even a small shift in a patient's description can lead to a dangerous misdiagnosis.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.