Performance, safety, and repeated-query stability of large language models versus early-career urologists in urological oncology: a source-masked vignette benchmark
This study demonstrates that while early-career urologists significantly outperform consumer-facing large language models in urological oncology regarding overall accuracy, safety, and response stability, the models' frequent safety-critical errors and inconsistent outputs underscore the necessity for clinician supervision rather than autonomous deployment.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the rapidly evolving landscape of modern medicine, artificial intelligence has begun to offer a new kind of assistance, capable of reading vast libraries of medical literature and answering complex questions in seconds. These systems, known as large language models, function by predicting the most likely words to follow a prompt, allowing them to generate coherent and often convincing responses to clinical scenarios. For doctors, this technology promises a powerful tool to organize information, check guidelines, and support decision-making. However, the medical field is built on a foundation of precision and safety, where a single error in judgment can have life-altering consequences. The critical question facing the medical community is not just whether these machines can sound knowledgeable, but whether they can be trusted to make safe, consistent, and complete decisions when a patient's health is on the line.
To answer this, a team of researchers in Turkey designed a rigorous test to see how well three popular consumer-facing artificial intelligence systems performed in the high-stakes world of urological oncology, which deals with cancers of the urinary system. They created one hundred detailed, synthetic patient stories, covering five different types of cancer and various stages of care, from initial diagnosis to long-term follow-up. These stories were presented to the three artificial intelligence models, as well as to two early-career urologists who had recently finished their specialized training. The doctors and the machines were given the exact same instructions and were not allowed to look up answers or use outside resources. Two independent experts, who did not know whether the answers came from a human or a machine, then scored every single response against a strict set of guidelines established by the European Association of Urology. The scoring looked for four things: whether the advice matched current medical guidelines, whether it was safe for the patient, whether it covered all necessary steps, and whether the disease stage was identified correctly.
The results revealed a clear gap between the performance of the human doctors and the artificial intelligence systems. The two early-career urologists achieved a high level of success, providing safe and complete answers in eighty-five and eighty-nine percent of the cases. In contrast, the artificial intelligence models were successful in only sixty to seventy percent of the cases. While the machines often produced answers that looked correct on the surface, they frequently missed crucial details or failed to include specific safety warnings that the human doctors caught. The most concerning finding was related to patient safety. Approximately one out of every four responses from the artificial intelligence systems contained an error that could have caused serious harm, such as recommending a treatment that was dangerous for the patient or missing a critical warning sign. The human doctors, by comparison, made such safety-critical errors in only one or two percent of the cases.
Beyond the accuracy of the answers, the study also examined how reliable the artificial intelligence systems were when asked the same question multiple times. In the real world, a doctor would give the same advice for the same patient every time they reviewed the case. The researchers found that the artificial intelligence models were surprisingly inconsistent. When the same patient story was asked three times in a row, the system gave the exact same classification of success or failure less than half the time. Even more troubling, the system sometimes gave a safe answer on the first try, an unsafe answer on the second, and a safe answer again on the third. This lack of stability means that the output of these tools can change unpredictably, making it difficult to rely on them for consistent medical advice.
The researchers concluded that while these artificial intelligence systems can generate useful information, they are not yet ready to operate independently in a clinical setting. The study suggests that the best use for these tools right now is as a supervised assistant, where a human doctor reviews every recommendation before it is applied to a patient. The machines can help organize information or suggest options, but the final decision must remain in human hands to ensure safety and consistency. Until these systems can match the reliability and safety record of human specialists, their role in cancer care should be supportive rather than autonomous, serving as a second pair of eyes rather than the primary decision-maker.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.