← Latest papers
💻 computer science

Performance evaluation and benchmarking across 16 large language models on a comprehensive real-world emergency department triage data set

This study benchmarks 16 large language models on real-world emergency department triage data, finding that while structured prompting can achieve substantial agreement with human nurses for severity classification, most models exhibit limited accuracy, poor sectoral assignment, systematic overconfidence, and non-deterministic behavior, indicating they are not yet ready for clinical implementation without further validation and improvements.

Original authors: Leo Benning, Anja Hirsch, Matthias Gröschel, Tobias Röschl, Martin Spott, Felix Patricius Hans, Tim Urban, Hans-Jörg Busch, Alexander Meyer, Julian Gabriel Madrid

Published 2026-06-28
📖 4 min read☕ Coffee break read

Original authors: Leo Benning, Anja Hirsch, Matthias Gröschel, Tobias Röschl, Martin Spott, Felix Patricius Hans, Tim Urban, Hans-Jörg Busch, Alexander Meyer, Julian Gabriel Madrid

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a busy hospital emergency room as a giant, chaotic train station. Every day, thousands of people arrive with different problems, from a scraped knee to a heart attack. The job of the "station master" (the triage nurse) is to look at each person, decide how urgent their problem is, and tell them which platform to go to: the high-speed express train (the Emergency Department for critical cases) or the local bus stop (an urgent care clinic for minor issues).

This job is incredibly hard. It happens fast, under pressure, and the consequences of a wrong decision can be life-or-death.

Recently, people have been asking: "Can we hire a super-smart AI robot (a Large Language Model or LLM) to help the station master?" These AI robots are like brilliant students who have read almost every book in the library. But have they actually learned how to run a train station?

This paper is the report card for 16 different AI robots tested on real-life data from a German hospital. Here is what they found, explained simply:

1. The "Test Drive"

The researchers took 16,107 real patient records (anonymized, so no names were involved) and asked 16 different AI models to act as the triage nurse. They had to do two things:

  • Grade the urgency: Give the patient a score from 1 (dying right now) to 5 (a minor inconvenience).
  • Pick the destination: Send them to the Emergency Room (ER) or the Urgent Care Clinic.

They compared the AI's answers to what the actual human nurses decided.

2. The Results: Most Robots Are Still Learning

If the human nurses were the "Gold Standard," most of the AI robots didn't pass the test with flying colors.

  • The "Average" Student: Most of the AI models got a grade that was only "fair" to "moderate." They agreed with the nurses about half the time.
  • The "Confident" Failure: Here is the scary part. Even when the AI was wrong, it was extremely confident about being right. It's like a student who gets a math problem wrong but raises their hand and shouts, "I'm 100% sure this is the answer!" The paper found that the AI's self-confidence had almost nothing to do with whether it was actually correct.
  • The "Unreliable" Robot: If you asked the same AI the same question 100 times, it gave a different answer 23% of the time. Imagine asking a GPS for directions, and it says "Turn Left" today, but "Turn Right" tomorrow, even though the traffic hasn't changed. This "non-deterministic" behavior is dangerous in a hospital.

3. The Secret Weapon: The "Step-by-Step" Guide

There was one big surprise. The researchers tried a specific trick with one of the best AI models. Instead of just saying, "Here is a patient, what do you think?", they gave the AI a checklist.

  • Step 1: Is the patient dying?
  • Step 2: Do they have high-risk symptoms?
  • Step 3: What are their vital signs?
  • Step 4: How many resources will they need?

When the AI was forced to follow this step-by-step logic (like a human nurse does), its performance jumped up to "substantial." It became almost as good as a human nurse.
The Lesson: It didn't matter how "big" or "smart" the AI was. What mattered was how they asked the question. Giving the AI a clear, structured rulebook worked much better than just letting it guess.

4. The "Where to Go" Problem

While the AI was okay at guessing the urgency score (sometimes), it was terrible at deciding where the patient should go (ER vs. Urgent Care).

  • The AI often sent people with minor problems to the ER (over-triage), which clogs up the system.
  • Or, it sent people with serious problems to the Urgent Care (under-triage), which is dangerous.
  • The paper suggests this is because the AI doesn't know the "local rules" of that specific hospital, like which clinic is open at 3 AM or which building has the right equipment.

5. The Bottom Line

The paper concludes that while these AI robots are impressive, they are not ready to drive the bus alone.

  • They are too inconsistent (they change their minds).
  • They are too overconfident (they don't know when they are wrong).
  • They need a human "co-pilot" to check their work.

The only way to make them work better right now is to stop treating them like magic black boxes and start treating them like students who need a strict, step-by-step instruction manual. Until we can fix their reliability and confidence issues, they should not be making life-or-death decisions on their own.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →