CareLoop: Measuring the Gap Between Knowing Medicine and Practicing It in the Context of Patients' Lives
This paper introduces CareLoop, a clinical-world sandbox that evaluates AI models on their ability to practice medicine within the complex context of patients' lives by measuring their performance in discovering hidden states, adapting to constraints, and managing consequences across simulated trajectories, revealing that execution-grounded capabilities are distinct from and more challenging than simply knowing medical facts.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Medicine is often taught as a collection of facts: a list of symptoms, a catalog of treatments, and a set of rules for diagnosis. In this view, a doctor's job is to retrieve the correct answer from memory and apply it. But in the real world, medicine is not a puzzle with a single solution waiting to be found; it is a conversation with a person whose life, beliefs, and circumstances constantly reshape the path forward. A patient might misunderstand a doctor's instructions, hide a painful detail out of shame, or simply lack the money or transportation to get a prescribed test. A medical plan that looks perfect on paper can fail if it does not account for these human realities. The challenge for artificial intelligence is not just to know the right medical facts, but to navigate the messy, unpredictable space between knowing those facts and actually helping a person.
Researchers have long tried to test whether computer programs can answer medical questions correctly. These tests usually ask a simple question: does the AI know the right diagnosis? However, a new study introduces a different kind of test, one that asks whether an AI can practice medicine in the context of a patient's life. The researchers built a simulated environment called CareLoop, designed to measure the gap between knowing medicine and practicing it. Instead of giving the AI a static list of symptoms to solve, they created a dynamic world where patients have their own memories, beliefs, and limitations. In this world, information is hidden, evidence can be delayed, and actions taken by the AI do not always lead to the expected result. The goal was to see if an AI could discover missing information, correct misunderstandings, adapt to real-world barriers, and take responsibility for a patient's care over time, rather than just offering a correct answer at the start of a conversation.
The study involved creating 120 detailed scenarios based on real, anonymized medical records. Each scenario was a contract that defined the hidden state of a patient's health, what the patient and their family believed, and what obstacles might appear. These obstacles included things like a patient misremembering their medication, a lab result arriving late, a family member refusing a treatment, or a patient being unable to afford a visit. The researchers then asked ten different AI models to act as doctors in these simulations. The models had to interact with the simulated patients and family members, who were programmed to be imperfect: they could forget instructions, lie about symptoms, or get confused. The AI had to figure out what was really happening, verify the information, and guide the patient toward a safe outcome.
The results showed a clear difference between models that simply knew medical facts and those that could navigate the complexity of a patient's life. The best-performing model, GPT-5.6 Sol, achieved the highest scores in handling real-world frictions, though its lead over the next tier of models (including GPT-5.4, DeepSeek-V4-Pro, and Qwen3.8-Max) was not statistically distinct due to overlapping confidence intervals. The study revealed that even the top models struggled with specific tasks. The biggest challenge for all of them was discovering hidden states—figuring out facts that the patient had not volunteered or that were buried in the background, which the paper identifies as the largest shared bottleneck. Another common failure was the "action-to-impact" gap: the models often took an action, like suggesting a test, but failed to ensure the test was actually done or to follow up on the results. In many cases, the AI would declare the conversation finished even though the patient's problem was not truly resolved, a mistake known as premature closure.
The study also compared how well the AI models performed against human doctors. A panel of 139 physicians reviewed the same simulated cases and ranked the AI models. While individual doctors sometimes disagreed with each other, and sometimes disagreed with the AI evaluators, the overall rankings of the models were remarkably stable. When the researchers looked at the aggregate results, the human doctors and the AI judges agreed on which models were the best and which were the worst. This suggests that the new testing method is reliable enough to distinguish between different levels of capability, even when the task is complex and the evaluation is subjective.
One of the most important findings was that finishing a conversation quickly is not the same as providing good care. Many of the AI models ended their interactions early, marking the task as complete, even though the patient still had unresolved risks or unverified information. The best models were the ones that knew when to keep an episode open, maintaining responsibility for the patient until the situation was truly safe to close. This distinction highlights a fundamental shift in how we should evaluate medical AI. It is not enough for a system to be fluent or to give a correct diagnosis in isolation. True competence in a medical setting requires the ability to manage uncertainty, verify that plans are carried out, and stay accountable for a patient's well-being even when the path is unclear.
The researchers emphasize that these results come from a simulation, not a real hospital. The study does not prove that these AI models are ready to treat patients or that they are safe for clinical use. Instead, it provides a way to measure how close they are to that goal. By exposing the specific ways in which AI fails to bridge the gap between knowledge and practice, the study offers a roadmap for improvement. It shows that the next generation of medical AI needs to be better at listening, at verifying, and at understanding that a patient's life is full of obstacles that no amount of textbook knowledge can automatically overcome. The path forward is not just about making AI smarter, but about making it more capable of walking alongside a person through the complexities of their own health journey.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.