Position: Medical AI Neglects Real Treatment Outcomes
This position paper argues that medical AI's reliance on human opinions and textual syntheses rather than actual treatment outcome data from observational and experimental sources severely limits its potential, necessitating a shift toward incorporating real-world outcome data into both training and evaluation to better achieve the ultimate goal of improving patient health.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Medicine has always been a field where the difference between a good guess and a life-saving decision can be a matter of data. For decades, doctors have moved away from relying solely on the intuition of individual experts, shifting instead toward a system built on the collective experience of thousands of patients. This approach, known as evidence-based medicine, seeks to ground every decision in the actual results of treatments rather than just the theories behind them. It is a system that asks not just what a drug is supposed to do, but what it actually did for the people who took it. In recent years, a new tool has entered this landscape: artificial intelligence. These computer systems, often called large language models, have shown a remarkable ability to read medical texts, diagnose conditions, and predict future health risks. They are trained on vast libraries of medical journals, textbooks, and clinical guidelines, learning to mimic the language and reasoning of human experts.
However, a critical gap has emerged in how these intelligent systems are built and tested. While they are becoming better at identifying diseases or predicting who might get sick, they are not being taught to understand the true consequences of the treatments they recommend. The core question facing researchers today is whether these machines are learning from the stories written in medical books, or from the hard, often messy reality of patient outcomes. If an AI is trained only on the opinions of experts and the rules written in guidelines, it may learn to sound like a doctor without ever truly understanding what happens when a treatment is applied to a real human being. This distinction is not just an academic detail; it is the difference between a system that recites rules and one that can genuinely improve health.
A new position paper argues that the current generation of medical artificial intelligence is failing to learn from real treatment outcomes. The authors, researchers from the Harvard Pilgrim Health Care Institute and Harvard Medical School, contend that while these models are being fed massive amounts of medical text, they are largely missing the most important data of all: the records of what actually happened to patients after they received care. The paper suggests that the field has become too focused on intermediate steps, such as diagnosing a condition or predicting a risk, while neglecting the ultimate goal of medicine: improving the patient's health through treatment. The researchers point out that most medical AI models are trained on static documents like clinical guidelines, drug labels, and published case reports. These sources represent the consensus of the medical community at a specific moment in time, but they do not capture the dynamic, longitudinal reality of how treatments play out over time in diverse populations.
To illustrate the danger of relying on these static sources, the researchers examined how current AI models handle complex medical scenarios. They looked at a specific case involving a patient with a rare form of pancreatitis who was successfully treated with a drug called ivacaftor. In the real world, this treatment worked, but the official label for the drug, which is a government-approved document used to train many AI systems, did not list this specific use. When the researchers asked an AI model to evaluate the utility of this drug for the patient, the model, relying on the official label, incorrectly concluded that the drug would be ineffective. The model was essentially following the rules in the book rather than the reality of the patient's recovery. This error occurred even though the AI had access to the case report describing the successful treatment; the model's training on the rigid drug label overrode the evidence of the actual outcome. The researchers found that this problem is widespread. When they tested other models against similar scenarios involving clinical guidelines, the systems often failed to recognize complex situations where a doctor might need to deviate from the standard rules to help a specific patient.
The paper further highlights that the way these AI systems are evaluated is part of the problem. Many benchmarks, or tests used to measure how well an AI is performing, rely on human experts to grade the answers. The researchers argue that asking a group of doctors to agree on an answer is not the same as knowing the truth. In one analysis, they found that a significant portion of the criteria used to grade AI answers were simply direct quotes from guidelines, rather than reflections of actual clinical success. This creates a circular logic where the AI is trained on guidelines, tested against guidelines, and then praised for following guidelines, even if those guidelines are outdated or do not apply to every individual. The researchers suggest that this approach limits the potential of artificial intelligence to become a true partner in medicine, capable of discovering new insights rather than just repeating what has already been written.
The authors propose a fundamental shift in how medical AI is developed. Instead of training these systems primarily on text, they recommend using longitudinal patient data, such as electronic health records and insurance claims, which track what treatments were given and what happened to the patients afterward. These records contain the raw material of real-world evidence, showing the cause-and-effect relationships between a therapy and a patient's recovery or decline. The researchers also suggest that the testing of these models should move away from asking them to recite guidelines and toward asking them to predict the outcomes of real-world experiments or to emulate the results of clinical trials. By grounding the training and evaluation of these systems in actual patient outcomes, the field could move toward a future where artificial intelligence helps doctors make decisions that are not just theoretically sound, but practically effective for the people sitting in front of them.
The paper does not suggest that current medical AI is useless or that the existing benchmarks are entirely without value. The researchers acknowledge that these tools have made rapid progress and have demonstrated valuable capabilities in areas like diagnosis and information retrieval. However, they argue that the field has reached a point where continuing to rely on the same methods of training and evaluation will prevent the next leap forward. The goal is not to discard the knowledge contained in medical literature, but to ensure that the artificial intelligence systems built to use that knowledge are also anchored in the reality of patient care. By incorporating real treatment outcomes into the core of their development, the researchers believe medical AI can evolve from a system that reads the rules to one that understands the consequences, ultimately helping to write better guidelines and improve the health of patients everywhere.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.