Automating the triage of rheumatology outpatient referrals: a comparative evaluation of 23 large language models under simple and advanced prompting
This study demonstrates that contemporary large language models, particularly when guided by advanced prompting techniques, can automate rheumatology referral triage with accuracy comparable to or exceeding human experts, effectively eliminating the need for expensive, large-scale models.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Every day, thousands of people with joint pain, swelling, or stiffness write to rheumatologists asking for help. These letters, known as referrals, arrive in a flood that specialists must sort through. The goal of this sorting, called triage, is to decide who needs to be seen immediately because their condition could threaten an organ or their life, and who can safely wait weeks or months. This task falls on senior doctors, who must read complex, unstructured text to make a judgment call. It is a high-volume, mentally exhausting job that takes time away from seeing patients and treating disease. When the system works well, the sickest people get help fast. When it fails, a patient with a rapidly worsening autoimmune condition might wait too long, risking permanent damage. Yet, even when done by experienced human doctors, this sorting process is not perfect; experts often disagree with one another, and urgent cases are sometimes missed or delayed.
As the number of referrals grows faster than the number of available specialists, clinics are looking for ways to use artificial intelligence to help. Specifically, they are testing large language models, which are computer programs trained on vast amounts of text that can understand and generate human language. These models are already being used to answer medical questions and assist with diagnosis, but it was unclear if they could reliably handle the specific, high-stakes task of sorting rheumatology referrals. The central question was not just whether a computer could do this, but whether it could do it as well as a human team, and whether the cost of using the most powerful, expensive computers was necessary to get the job done right.
A researcher at Monash Health in Australia set out to answer these questions by putting twenty-three different large language models through a rigorous test. They created twenty realistic referral scenarios based on real cases, covering the full range of urgency from life-threatening emergencies to routine aches. To establish a gold standard for the correct answer, four independent rheumatologists reviewed every single case without knowing what the others decided, eventually agreeing on a consensus urgency level for each. The researcher then asked each of the twenty-three computer models to triage these same twenty cases. To ensure the results were stable and not just lucky guesses, each model was asked to make the decision three times for every case. The experiment was run twice: once with a very simple instruction telling the model to sort the referrals, and again with a much more detailed instruction that explained exactly how to think about the problem, provided examples of what each urgency level looked like, and guided the model to focus on specific medical signs while ignoring less relevant factors like pain levels.
The results revealed a clear pattern in how these machines think. When given only a simple instruction, the models performed very differently from one another. The largest, most advanced, and most expensive models were significantly more accurate than the smaller, cheaper ones. In this simple setting, paying more for a bigger model bought better performance. However, the story changed completely when the researcher switched to the advanced, detailed instructions. With this better guidance, the performance gap between the models vanished. The smaller, cheaper models caught up to the expensive ones, and the link between cost and accuracy disappeared. The detailed instructions essentially taught the smaller models how to reason through the problem, allowing them to perform just as well as the giants without the high price tag.
Even with the best instructions, the computers were not flawless. They still made mistakes, sometimes suggesting a patient was sicker than they were, and sometimes suggesting they were less sick. The researcher noted that the error of under-triaging—missing an urgent case—was the most dangerous type of mistake, as it could delay critical care. Despite this risk, the top-performing models matched the human consensus on most cases, often agreeing on the correct urgency level in all three attempts. This level of performance placed them within or above the range of accuracy reported for human doctors, who are known to disagree with each other on about a third of cases and frequently change their urgency rating after seeing the patient.
The study suggests that the tools to automate this administrative burden already exist and may be more affordable than previously thought. The key finding is that the way a question is asked to the computer matters as much as the computer itself. By using a well-designed prompt that includes clear rules and examples, clinics might be able to use smaller, cheaper models to achieve results that rival the most expensive systems. This could allow medical services to return valuable time to senior doctors, freeing them from sorting letters so they can focus on treating patients. While the technology is promising, the researcher emphasizes that any system used in the real world must be carefully monitored to ensure it does not miss urgent cases, and that human doctors must remain in the loop to review the most critical decisions. The path forward involves testing these tools on real, live referrals and refining the instructions to ensure safety, but the data indicates that a reliable, automated assistant for rheumatology triage is no longer a distant dream.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.