Evaluating Safety-Critical Communication Behavior of Large Language Models Using Workflow-Embedded Multi-Agent Clinical Simulation
This study demonstrates that embedding large language models in a role-locked, multi-agent clinical simulation reveals that while structured ISBAR handovers reduce hallucinations and improve communication efficiency, they fail to eliminate safety-critical omissions, highlighting the continued necessity of expert human oversight over automated evaluation systems for assessing clinical safety.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Hospitals rely on a delicate chain of communication to keep patients safe. When a nurse notices a patient getting worse, they must quickly call a senior doctor, known as a registrar, to report the problem and get instructions. This moment of escalation is critical; if the nurse forgets a key detail, or if the doctor misunderstands the urgency, a patient could suffer preventable harm. To help, many hospitals use a standard checklist called ISBAR, which stands for Identify, Situation, Background, Assessment, and Recommendation. This tool forces the speaker to organize their thoughts before speaking, ensuring that vital information like vital signs and specific requests are not left out.
As hospitals begin to explore using artificial intelligence to assist with these conversations, a new question arises: can a computer program handle this high-stakes communication safely? Large language models, the technology behind modern chatbots, are excellent at generating human-like text. However, they are also known to sometimes invent facts or miss crucial details when left to their own devices. Before any such system could be trusted in a real hospital, researchers needed a way to test if these programs would behave safely when placed inside a realistic, time-pressured medical workflow, rather than just answering simple questions on a screen.
To answer this, a team of researchers at University College Cork built a digital simulation that mimics the exact flow of a hospital ward. They did not use real patients or real doctors. Instead, they created a controlled environment where artificial intelligence agents played specific roles: one agent acted as the ward nurse, another as the senior registrar, and a third as a nurse manager. These agents were "role-locked," meaning the computer programs were strictly instructed to stay in character, never breaking their assigned job to narrate the story or act like a different person. The researchers fed these agents synthetic case files describing patients who were getting sick, such as someone with a severe infection or a bleeding wound. The goal was to watch how the AI agents talked to each other when the patient's condition worsened.
The researchers set up a fair test by running the same scenario twice for each patient case. In the first version, the nurse agent started the conversation in a natural, unstructured way, just as a human might speak without a script. In the second version, the nurse was required to use the ISBAR checklist, delivering the information in the specific, organized format. The registrar agent then responded to both types of calls, and the researchers recorded every word of the resulting dialogue. They then used a sophisticated automated system to read through hundreds of these conversations, looking for specific types of errors. They checked to see if the nurse left out any life-saving details, if the doctor gave a clear plan with a specific time for follow-up, and if the AI made up facts that were not in the patient's file.
The results revealed a surprising mix of success and failure. When the nurse used the ISBAR structure, the conversations became much more efficient. The dialogue was shorter, taking fewer turns to reach a decision, and the AI was far less likely to invent fake medical facts. In the unstructured conversations, the AI made up details about the patient's condition in about 26.5 percent of the cases, but when the checklist was used, that number dropped to 10.0 percent. This suggests that giving the AI a strict format helps it stay grounded in the facts provided to it.
However, the study also uncovered a hidden danger. While the structured conversations were cleaner and faster, they did not make the communication safer in the most critical way. The researchers found that the rate of missing vital information, known as safety-critical omissions, did not go down. In fact, the structured conversations had a slightly higher rate of these dangerous gaps, at 16.0 percent, compared to 11.5 percent in the unstructured ones. The researchers suspect that when the AI follows a rigid checklist, it may assume it has covered everything and stop asking the clarifying questions that a human might ask to fill in the blanks. The checklist prevented the AI from making things up, but it did not stop it from forgetting to say something essential.
To ensure their computer analysis was accurate, the researchers asked two experienced intensive care nurses to review a smaller selection of these conversations. The human experts looked for the same things the computer did: missing details, made-up facts, and clear action plans. The human nurses and the automated system did not always agree. The computer was very good at spotting potential missing information, but it was often too strict, flagging omissions that the human nurses considered unimportant. Conversely, the computer sometimes missed the subtle ways in which a plan lacked real-world safety. This gap showed that while computers can scan thousands of conversations quickly, they still do not fully understand the nuance of clinical safety the way a trained human does.
Ultimately, this work demonstrates that testing artificial intelligence in a realistic, role-playing simulation is essential for understanding how it will behave in the real world. The study showed that simply making a system follow a communication checklist improves its efficiency and reduces its tendency to lie, but it does not guarantee that it will not miss critical safety details. The researchers concluded that before such tools are used in hospitals, they must be evaluated not just on whether they sound polite or follow a format, but on whether they can reliably convey the complete picture of a patient's condition. The path to safe medical AI requires more than just better software; it requires a rigorous testing process that combines automated checks with the irreplaceable judgment of human experts.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.