Concordance of a ChatGPT-4o–based triage model with emergency physician categorization in the emergency department: a prospective observational comparison with emergency medical technicians
This prospective observational study found that a ChatGPT-4o-based triage model demonstrated significantly higher concordance with emergency physician categorization than emergency medical technicians when using identical structured clinical data, though this agreement reflects rater consistency rather than diagnostic validity and is accompanied by a tendency to over-triage high-acuity patients.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a busy emergency room as a massive, chaotic train station. Every minute, new passengers (patients) arrive, and someone needs to decide who gets on the first train (immediate care), who waits on the platform (moderate care), and who can walk to the ticket counter later (low urgency). This process is called triage.
Usually, a human expert (an Emergency Physician) makes these decisions. But humans get tired, distracted, or influenced by how they're feeling, so two different experts might sort the same passenger differently.
This study asked a simple question: Can a super-smart computer program (ChatGPT-4o) sort these passengers just as well as the human expert?
Here is the breakdown of what the researchers did and found, using simple analogies:
The Setup: The "Sorting Game"
The researchers set up a game with three players, all looking at the same 788 passengers arriving at a Turkish hospital:
- The Human Expert (The Gold Standard): A single, highly trained Emergency Doctor. This doctor is the "referee."
- The Computer (The New Player): A ChatGPT-4o model. It didn't see the patients in person; it only read a standardized checklist of facts about them (symptoms, history, etc.).
- The Triage Tech (The Real-World Player): Emergency Medical Technicians (EMTs). These are the people who usually do the sorting at the front door in this specific hospital. They saw the patients in person, just like in real life.
The Rules:
- The Doctor and the Computer both looked at the exact same written checklist for every patient.
- The EMTs looked at the actual people standing in front of them.
- Everyone had to put the patient into one of three colored bins: Red (Emergency!), Yellow (Wait a bit), or Green (Fine, go later).
The Results: Who Won the Game?
1. The Computer vs. The Doctor
The computer was incredibly good at mimicking the doctor.
- The Score: If you imagine a scale where 1.0 is perfect agreement, the computer and the doctor agreed about 84% of the time (a score of 0.84).
- The Metaphor: It was like a student who studied the teacher's answer key so thoroughly that they wrote the exact same answers on the test. They were almost identical in how they sorted the patients.
2. The Computer vs. The Triage Tech
The computer also did better than the human techs who were actually standing at the door.
- The Score: The techs agreed with the doctor about 63% of the time.
- The Gap: The computer was significantly better at matching the doctor's choices than the techs were.
3. The "Red" Bin Problem (The Catch)
Here is where it gets tricky. The "Red" bin is for life-threatening emergencies. Missing a Red patient is dangerous.
- The Techs: They were very careful. If they put someone in Red, they were usually right (87% accuracy). However, they missed a lot of actual emergencies, putting 45 out of 79 critical patients into the "Yellow" bin instead. They were under-triaging (playing it too safe).
- The Computer: It caught almost all the critical patients (92% of the "Red" patients were correctly identified). But, it was also very "paranoid." It put 28 people who were actually "Yellow" into the "Red" bin.
- The Metaphor: The computer was like a security guard who screams "FIRE!" at the slightest smell of smoke. It saves everyone who is actually on fire, but it also causes a lot of false alarms, clogging up the emergency exits with people who didn't need them.
Important Warnings (The Fine Print)
The authors are very careful not to say the computer is "better" than a doctor in real life. Here is why:
- The "Shared Textbook" Bias: The computer and the doctor were both looking at the same written notes. The techs were looking at the people. Because the computer and doctor were working from the same source material, they were more likely to agree. It's like two people grading a test who both have the answer key; they will agree more than someone who has to guess based on a student's messy handwriting.
- No "Real Life" Test: The study didn't check if the patients actually got better or worse. It only checked if the computer agreed with the doctor's opinion. We don't know if the computer's "Red" calls actually saved lives or just wasted resources.
- Missing the Worst Cases: The study excluded the sickest patients (those who were unconscious or needed CPR immediately). So, we don't know how the computer handles the absolute worst emergencies.
- One Doctor: The "Gold Standard" was just one doctor's opinion. If that doctor had a bad day or a specific bias, the computer just copied that bias.
The Bottom Line
This study is a proof-of-concept. It's like showing that a new type of engine can run on a test track at high speeds.
- What it proved: A ChatGPT-4o model can read a patient's chart and sort them into categories almost exactly like a specialist doctor does, and it does it more consistently than the front-line techs in this specific setting.
- What it didn't prove: It didn't prove the computer is safe to use in a real hospital yet. It tends to over-call emergencies (false alarms), and we don't know if it works on the sickest patients or with different languages and hospital systems.
The researchers conclude that this technology is promising but needs much more testing before it can replace or even assist humans in a real emergency room.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.