Development and Validation of a Localized Large Language Model for Preoperative Anesthesia Planning: Protocol for a Retrospective Observational Study and Clinical Turing Test
This paper outlines a protocol for developing and validating a localized Large Language Model, fine-tuned on 3,000 historical anesthesia records from Istishari Hospital, to generate preoperative anesthesia plans that are evaluated against human expert standards for safety and accuracy in a retrospective "Clinical Turing Test."
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Before a patient enters an operating room, an anesthesiologist must perform a complex mental task: they take a person's medical history, their current vital signs, and the specific demands of the upcoming surgery, then weave these facts into a single, safe plan for keeping the patient asleep and pain-free. This process requires rapid thinking and deep knowledge, but it is also prone to human error, especially when doctors are tired or when different doctors interpret the same information in slightly different ways. In recent years, computers have begun to help with medical tasks, mostly by predicting things like blood pressure drops or managing drug pumps automatically. However, a new type of computer program, known as a large language model, has emerged with a different ability. These programs can read vast amounts of text and write new text that sounds like a human expert. While these tools are already used for studying and testing, no one has yet proven they can safely write a real, customized medical plan for a specific hospital, using that hospital's own rules and habits, without making dangerous mistakes.
A researcher at Istishari Hospital in Amman, Jordan, set out to test this idea. They wanted to see if they could teach a computer to act like a senior anesthesiologist by feeding it thousands of real, past cases from their own hospital. Their goal was to create a "digital fellow," a specialized computer program that could draft a complete anesthesia plan just as a human would. To do this, they gathered 3,000 historical records of patients who had undergone surgery. These records were not digital files but physical paper documents containing notes on the patient's height, weight, heart history, lung conditions, and the exact drugs and airway tools the human team had used during the operation. A team of junior doctors and interns manually read through every single paper record, copying the information into a secure digital database. They were careful to remove all names and identifying details, leaving only the medical facts and the resulting treatment plans.
Once this massive collection of data was ready, the researcher used it to train a powerful computer model. They instructed the model to learn the patterns of how their specific hospital handles anesthesia. The computer was told to act as a senior consultant at that hospital and to write a new plan for any patient based on the data it had learned. The plan had to cover four main areas: how to prepare the patient before surgery, how to put them to sleep and manage their breathing, how to keep them stable during the procedure, and how to wake them up and manage their pain afterward. The researcher did not just let the computer guess; they built the system on the actual, successful decisions made by human doctors in their own institution, hoping this would stop the computer from inventing facts or suggesting treatments that do not exist in their hospital.
To see if the computer's work was any good, the researcher organized a blind test, similar to a challenge where a person tries to guess which of two writers is human and which is a machine. They took a set of patient cases and showed them to a group of senior anesthesiologists. For each case, the experts were given two different plans: one was the original plan written by a human doctor years ago, and the other was the new plan written by the computer. The experts did not know which was which. They were asked to judge the plans based on safety, checking for any dangerous errors, and on completeness, ensuring all necessary steps were included. Finally, they had to guess which plan was written by the computer.
The study is currently a protocol, meaning it is the detailed plan for how the research will be done, rather than a report of final results. The researcher proposes that if their computer model is trained on these 3,000 real local cases, it will produce plans that are just as safe and accurate as those written by human experts. They believe this approach will reduce the risk of the computer making up information, a problem known as "hallucination," because the computer is grounded in the reality of their own hospital's practices. If the test shows that the experts cannot tell the difference between the human and computer plans, it suggests that this technology could one day help doctors by handling the heavy mental work of writing routine plans, allowing them to focus more on the patient and less on paperwork. The project also aims to turn the difficult task of reading thousands of old paper records into a learning experience for the junior doctors who do the data entry, forcing them to study the decisions of their mentors while they build the database.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.