← Latest papers
💬 NLP

ResidencyRL: Reinforcement Learning in Simulated Clinical Environments

ResidencyRL introduces a reinforcement learning framework that trains clinical AI agents through simulated multi-turn patient encounters with adversarial LLMs and structured rewards, significantly improving diagnostic accuracy, safety, and communication skills while demonstrating robust generalization across diverse clinical benchmarks.

Original authors: Valentin Liévin, Samuel Schmidgall, Tim Strother, Alex Bijamov, Akshay Goel, Anil Palepu, Chunjong Park, Vahid Balazadeh, Min Woo Sun, Marius Guerard, Justin Chen, Dave Steiner, Vikram Dhillon, Ibrahi
Published 2026-08-10
📖 4 min read☕ Coffee break read

Original authors: Valentin Liévin, Samuel Schmidgall, Tim Strother, Alex Bijamov, Akshay Goel, Anil Palepu, Chunjong Park, Vahid Balazadeh, Min Woo Sun, Marius Guerard, Justin Chen, Dave Steiner, Vikram Dhillon, Ibrahim Azar, Akhil Mehta, Nicholas Spetsieris, Shilpan Shah, Maen Abdelrahim, Amit Dahiya, Yun Liu, Katherine Chou, Yossi Matias, Avinatan Hassidim, Dale R. Webster, Quoc V. Le, Raia Hadsell, Joelle Barral, Carey Radebaugh, Aleksandra Faust, Shekoofeh Azizi, Mike Schaekermann, Po-Hsuan Cameron Chen, Tao Tu, David Racz, Lin Yang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where learning to be a doctor isn't just about memorizing a giant library of medical facts, but about actually doing the job. In the real world, medical students start with textbooks, but they only become true experts after years of "residency"—a grueling apprenticeship where they talk to thousands of real patients, make mistakes, get feedback, and learn how to handle the messy, unpredictable nature of human illness. This is the corner of science known as Artificial Intelligence (AI) in healthcare. For a long time, AI was like a brilliant student who could ace every written test but froze when asked to talk to a real person. The key concepts here are Large Language Models (LLMs), which are super-smart computer programs trained on vast amounts of text, and Reinforcement Learning (RL), a training method where an AI learns by trying things, getting rewards for good moves, and penalties for bad ones, much like a dog learning tricks or a gamer mastering a video game. The big question everyone is asking is: Can we teach an AI to not just know medicine, but to practice it safely and effectively, just like a human resident does?

Enter ResidencyRL, a new experiment by a team of researchers who decided to build a digital "residency program" for an AI doctor. Instead of just feeding the AI more medical textbooks, they created a simulated hospital where the AI could practice talking to virtual patients. These weren't just simple chatbots; the virtual patients were designed to be tricky, sometimes hiding important details, acting confused, or even trying to trick the AI into making a mistake. The AI had to navigate long conversations (up to 60 turns!), ask the right questions, figure out what was wrong, and decide on a treatment plan, all while being graded by a strict digital supervisor.

The results of this digital residency were surprisingly promising. After training, the AI didn't just get better at the specific practice cases; it became a more thorough and careful doctor overall. In tests where the virtual patients were being particularly difficult or deceptive, the trained AI improved its diagnostic accuracy by 7.0% (jumping from 81.0% to 88.0%). More importantly, it stopped making a dangerous mistake called "premature closure"—where a doctor guesses the answer too quickly and stops looking for clues. The trained AI missed fewer critical warning signs, dropping its "missed red flag" rate by 31%. When real human doctors (who didn't know which AI was which) compared the trained AI against the untrained version, they preferred the trained one in 87.6% of the cases, praising it for gathering more complete information and creating safer, more sensible treatment plans.

The researchers found that these skills were like a universal toolkit. Even when they tested the AI on medical cases it had never seen before—like complex cancer scenarios or multi-visit chronic disease management—it still performed better than the untrained version. It suggests that the AI learned how to think like a doctor: how to be curious, how to dig deeper when things don't add up, and how to stay calm under pressure. However, the authors are careful to note that this is all happening in a simulation. While the AI is getting much better at the process of clinical reasoning, the team emphasizes that we still need to test it with real patients in the real world to be sure it can save lives without causing harm. For now, ResidencyRL shows that if you give an AI a chance to "practice" in a safe, simulated environment, it can learn to be a much more competent and reliable medical partner.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →