💻 computer science
Performance evaluation and benchmarking across 16 large language models on a comprehensive real-world emergency department triage data set
This study benchmarks 16 large language models on real-world emergency department triage data, finding that while structured prompting can achieve substantial agreement with human nurses for severity classification, most models exhibit limited accuracy, poor sectoral assignment, systematic overconfidence, and non-deterministic behavior, indicating they are not yet ready for clinical implementation without further validation and improvements.
Leo Benning, Anja Hirsch, Matthias Gröschel, Tobias Röschl, Martin Spott, Felix Patricius Hans, Tim Urban, Hans-Jörg Bus (…)2026-06-28