Developing an open-source framework for LLM evaluation of patients using EHR clinical documentation; performance of LLMs relative to medical professionals
This study demonstrates that current large language models fail to achieve inter-rater reliability comparable to medical professionals when extracting structured clinical information from ENT electronic health records, suggesting they are best suited for initial data extraction requiring human verification rather than autonomous clinical deployment.