Rare Diseases, Common Dilemmas: LLMs Prioritize Equal Resource Distribution over Patient Benefit in Decision-Making
This study reveals that state-of-the-art large language models, when presented with high-stakes rare disease scenarios, consistently prioritize equal resource distribution (justice) over patient-specific benefits (beneficence) and autonomy, a bias that is further modulated by whether decisions are framed as institutional or individual choices.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the high-stakes world of medicine, doctors often face choices where there is no single correct answer. When a patient is critically ill or facing a rare condition, the medical team must balance several competing moral duties. They must try to help the patient, avoid causing harm, respect the patient's right to choose their own path, and ensure that resources are shared fairly among everyone who needs them. These duties, known as beneficence, non-maleficence, autonomy, and justice, frequently pull in different directions. For instance, a treatment that might save a life could also cause immense suffering, or a scarce, expensive medicine might help one person deeply while leaving many others without care. As artificial intelligence becomes more common in healthcare, people are beginning to ask how these computer systems handle such difficult, value-laden decisions. The concern is not just whether the machine knows the facts, but whether it understands the weight of the moral trade-offs involved.
A new study investigates how large language models, the powerful AI systems capable of generating human-like text, make these ethical choices when caring for patients with rare diseases. Rare diseases present a unique challenge because they affect small numbers of people, often with limited medical data available. In these situations, families and doctors frequently have to make decisions under deep uncertainty, weighing imperfect options against each other. The researchers created a test to see if AI could navigate these murky waters. They built a collection of 208 realistic medical stories, or vignettes, based on real-world rare diseases and their specific genetic and clinical details. Each story presented a genuine dilemma where two different next steps were both medically reasonable but ethically opposed. For example, one option might prioritize the immediate needs of a single patient, while the other focused on the fair distribution of a limited resource across a group. The researchers then asked eleven different state-of-the-art AI models to choose between these two options.
The results revealed a striking pattern. Across all the different AI models tested, regardless of who built them or how they were trained, the systems consistently chose the option that prioritized justice. In the context of these medical stories, this meant the AI almost always favored equal distribution of resources over the specific, urgent needs of an individual patient. The models seemed to treat fairness as a matter of giving everyone the same share, rather than giving more to those who are sicker or have a greater medical necessity. This preference for equality appeared to be a superficial response to the concept of fairness, one that overlooked the severity of the patient's condition. The study suggests that these AI systems are not deeply reasoning through the moral complexity of the situation but are instead defaulting to a simple rule of equal distribution.
The researchers also discovered that the AI's choice depended heavily on who was described as the person making the final decision. When the story framed the decision as belonging to a committee or a hospital board, the AI strongly favored justice and equal resource sharing. However, when the decision was framed as belonging to a single doctor or the patient themselves, the AI shifted its preference. In those cases, the models moved away from strict equality and began to prioritize the patient's own wishes or their immediate well-being. This indicates that the AI is highly sensitive to the social role assigned to the decision-maker, mirroring human intuitions about how authority works. It appears to have learned that committees are responsible for fairness, while individuals are responsible for care and choice.
This behavior raises important questions about how these tools might be used in real hospitals. The study suggests that AI decision support systems might silently reinforce existing power structures rather than offering a fresh, principled perspective. If a doctor asks an AI for advice on a difficult case, the answer the AI gives could change simply based on whether the doctor asks as an individual or as part of a team. The researchers found that the models did not seem to be reasoning from first principles about what is best for the patient; instead, they were matching the situation to a pattern of expected behavior based on the role involved. This "authority bias" means that the AI might fail to challenge unfair institutional practices or to recognize when a patient's specific needs should override a general rule of equality.
The study does not claim that these AI models are broken or that they cannot be useful. Rather, it highlights a blind spot in how we currently evaluate medical AI. Most tests focus on whether the machine knows the right facts or can answer medical questions correctly. This research shows that even when models are factually accurate, they may still have hidden biases in how they weigh moral values. The findings suggest that as we move toward using AI in high-stakes medical settings, we must look beyond simple accuracy. We need to understand how these systems prioritize values and whether their choices align with the complex, nuanced reasoning that human doctors and ethicists use when lives are on the line. The study concludes that without careful attention to these ethical preferences, AI tools might end up simply repeating the status quo of institutional decision-making, potentially at the expense of the most vulnerable patients.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.