← Latest papers
📄 medicine

Inconsistent and divergent large language model safety guidance during intraoperative crises: a guideline- anchored benchmark

This study reveals that frontier large language models exhibit significant inconsistency and divergence in providing guideline-anchored safety guidance during simulated intraoperative crises, particularly regarding communication and escalation, highlighting the critical need to evaluate their reliability and reproducibility rather than aggregate accuracy alone for clinical decision support.

Original authors: Brian H. Park, Claire S. Soria, Preetham Suresh, Ryan Broderick, Bryan Sandler

Published 2026-09-07
📖 4 min read☕ Coffee break read

Original authors: Brian H. Park, Claire S. Soria, Preetham Suresh, Ryan Broderick, Bryan Sandler

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the high-stakes environment of an operating room, seconds matter. When a patient's condition suddenly worsens during surgery, the medical team must act immediately, following strict, life-saving checklists that have been refined over decades. These protocols are designed to catch errors before they cause harm, ensuring that the team communicates clearly, escalates the problem to the right people, and performs specific, critical maneuvers. As artificial intelligence tools known as large language models begin to enter the medical field, there is a growing hope that they could serve as a digital assistant, offering real-time guidance during these chaotic moments. However, unlike a human doctor who relies on training and experience, these computer programs generate answers based on patterns in data, and they can sometimes produce different responses to the exact same question. The critical question for patient safety is not just whether an artificial intelligence can identify a problem, but whether it can consistently recommend the correct, safe actions every single time a crisis occurs, and whether it can be trusted to do so without hesitation or contradiction.

A team of researchers from the University of California San Diego set out to test exactly this reliability. They created a rigorous simulation involving ten different types of surgical emergencies, ranging from severe bleeding to airway blockages and heart complications. Instead of asking the artificial intelligence to simply diagnose the problem, the researchers presented the models with a step-by-step unfolding crisis, asking after each stage what the surgical team should do next. The models were tested against a strict list of 150 specific safety actions derived from established medical guidelines and emergency manuals. These actions included things like declaring a crisis, calling for help, stopping a harmful procedure, or performing a life-saving maneuver. The researchers ran each scenario three times with three different leading artificial intelligence models, treating the computer programs as if they were new team members joining the operating room. The goal was to see if the models would give the same safe advice every time, or if their guidance would drift and change unpredictably.

The results revealed a significant and troubling inconsistency. When the same crisis scenario was presented to the same model multiple times, the computer did not give the same answer. In 29 out of the 30 combinations of model and scenario, the proportion of correct safety actions the model identified changed from one run to the next. On average, the difference in performance between two runs of the same model was substantial, varying by nearly 15 to 18 percentage points. This means that if a surgeon asked the same question twice during a real emergency, the computer might suggest a completely different set of actions the second time, potentially missing a critical step the first time around. While one model, Claude Opus 4.6, did identify more correct actions overall than the others, it still suffered from this same instability, and the differences in total accuracy between the models were not large enough to declare a clear winner. The most dangerous variations occurred in the area of communication and escalation. The models frequently failed to suggest that the team should declare an emergency or call for additional help, and the rate at which they missed these crucial steps varied wildly, with one model missing them nearly three times as often as another under identical conditions.

The study also examined whether the models would suggest actions that could actually harm the patient. Fortunately, these unsafe recommendations were rare across all the models tested. However, the fact that the models were inconsistent in their safe recommendations is a major concern. The researchers found that the models were generally better at suggesting specific medical procedures, like administering a drug or performing a surgery, but they struggled significantly with the softer, team-based aspects of crisis management, such as organizing the team or communicating the severity of the situation. This suggests that while the artificial intelligence might know the medical facts, it does not reliably understand the human dynamics required to manage a crisis safely. The researchers concluded that for these tools to be useful in an operating room, they must be judged not just on whether they are usually correct, but on whether they are consistently reproducible. A tool that gives different safety advice for the same problem cannot be trusted to support a medical team when a patient's life is on the line. The study serves as a benchmark, a measuring stick for future artificial intelligence, proving that reliability is just as important as accuracy when these systems are used in high-pressure medical situations.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →