← Latest papers
📄 otolaryngology

Developing an open-source framework for LLM evaluation of patients using EHR clinical documentation; performance of LLMs relative to medical professionals

This study demonstrates that current large language models fail to achieve inter-rater reliability comparable to medical professionals when extracting structured clinical information from ENT electronic health records, suggesting they are best suited for initial data extraction requiring human verification rather than autonomous clinical deployment.

Original authors: Barrett, L., Joshi, N., North, A. S., Dimitrov, L., Maughan, E. F., Ross, T., Pankhania, R., Paramjothy, K., Minty, I., Farache-Trajano, L., Smith, S. L., Mason, K. A., Bhargava, E. K., Donnelly, C.
Published 2026-08-24
📖 4 min read☕ Coffee break read

Original authors: Barrett, L., Joshi, N., North, A. S., Dimitrov, L., Maughan, E. F., Ross, T., Pankhania, R., Paramjothy, K., Minty, I., Farache-Trajano, L., Smith, S. L., Mason, K. A., Bhargava, E. K., Donnelly, C., Fatoum, H., Padiyar, A., Kader, Z., Chan, C. H. K., Schilder, A. G., Mehta, N.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Hospitals today generate a vast ocean of written records. Every patient visit, every test result, and every treatment decision is captured in digital notes known as electronic health records. These documents are rich with detail, containing the specific language doctors use to describe symptoms, diagnose conditions, and plan care. For decades, turning these unstructured stories into organized, searchable data has been a difficult task, often requiring humans to read and re-write the information by hand. Recently, a new type of computer program called a large language model has emerged. These systems are trained on massive amounts of text and can read, understand, and summarize human language with surprising fluency. They have shown they can answer medical questions and pass licensing exams, leading many to wonder if they can also read patient records and extract the vital facts inside them automatically.

This question is not just about saving time; it is about safety and accuracy. If a computer is to help doctors manage patient care, it must find the right information without making things up or missing critical details. A new study set out to test this capability directly. Researchers gathered nearly one hundred real patient records from ear, nose, and throat clinics and asked both human doctors and seven different large language models to read them. The task was to pull out specific facts, such as the patient's age, their symptoms, their diagnosis, and the treatments they received. To ensure everyone was speaking the same language, the researchers required that every piece of information be translated into a standardized code used by hospitals worldwide. This allowed for a precise comparison of how well the machines performed against the humans.

The results revealed a clear gap between human expertise and current artificial intelligence. When the fourteen doctors in the study read the same records, they agreed with each other on the extracted information about 81 percent of the time. This high level of agreement showed that the doctors were interpreting the notes consistently. However, when the computer models were compared to the doctors, the agreement dropped significantly to about 66 percent. In other words, the machines were not yet as reliable as the humans in identifying the same facts from the same text. The study tested models of different sizes, from massive systems with hundreds of billions of connections to smaller ones that could run on local computers. None of them reached the level of consistency shown by the medical professionals.

The performance varied depending on what kind of information was being extracted. The models were quite good at finding treatments and test results, which are often written in clear, standard ways. They were less successful with signs and diagnoses, which often require a deeper understanding of the clinical context. Interestingly, the computers tended to find more information than the doctors did, particularly regarding risk factors. While doctors often focus on the immediate problem, the models seemed to flag every possible risk mentioned in the text. This suggests the machines are not just copying human behavior but are processing the text differently, sometimes catching details a human might overlook, but also potentially flagging things that are not clinically relevant.

Despite these differences, the machines showed they could be useful tools if used correctly. The study found that when the models did identify a piece of information, they were usually right, with a precision rate of 97 percent. This means they rarely made up facts that were not there. However, they did miss some information about 15 percent of the time. This pattern points to a specific role for these tools in the future: they are best suited to act as a first pass, quickly scanning records to pull out likely candidates for the doctor to review. The final decision and verification would still need to come from a human. The researchers concluded that while these models are powerful, they are not yet ready to work alone in a clinical setting. They are best viewed as assistants that can help organize the flood of medical data, provided a human expert is always there to check their work.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →