← Latest papers
🤖 AI

High-Stakes Decisions with Language Models: Insights from Emergency Triage

This paper demonstrates that effectively deploying language models for high-stakes clinical decisions, such as emergency triage, requires integrating their predictive capabilities with explicit utility functions to weigh the consequences of different actions, rather than relying on accurate predictions alone.

Original authors: Khurram Yamin, Christopher Kelly, Bryan Wilder, Eric Horvitz

Published 2026-08-04
📖 4 min read☕ Coffee break read

Original authors: Khurram Yamin, Christopher Kelly, Bryan Wilder, Eric Horvitz

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are the captain of a spaceship navigating a foggy asteroid field. Your ship's computer is incredibly smart; it can look at a blurry dot and tell you, "There is a 90% chance that dot is a dangerous asteroid." But the computer doesn't know what to do with that number. If you crash, the ship explodes (a missed emergency). If you fire a warning shot at a harmless cloud, you waste fuel and annoy the crew (an unnecessary referral). The computer needs a rule: "If the danger is above 10%, fire the shot!" or "Only fire if it's above 90%!" This rule is called a utility function. In the real world, doctors face this exact foggy situation every day. They use triage to decide who needs immediate help and who can wait. For a long time, scientists have known that making these high-stakes calls isn't just about guessing the odds correctly; it's about knowing how much you hate making the wrong kind of mistake. If you are too scared of missing a crisis, you will send everyone to the hospital. If you are too worried about wasting time, you might miss a life-or-death situation. The big question is: when we ask super-smart computer brains (called Language Models) to be the captain, do they know how to set their own rules, or do they just guess?

This paper dives into that question by treating emergency medical triage not as a simple guessing game, but as a decision puzzle. The researchers took a set of medical stories (vignettes) where doctors had already decided who needed the ER and who didn't. They asked several different AI models to read these stories and first, guess the probability that a patient needed emergency care, and second, decide whether to send them to the ER. The team found something surprising: the AI models were actually really good at guessing the odds. They could tell the difference between a sick patient and a healthy one with high accuracy. However, when it came to making the final decision, the models were acting like they had a secret, unspoken rule that was way too cautious about sending people to the hospital.

The authors discovered that the "bad" behavior seen in a popular consumer health tool (which was found to miss more than half of real emergencies) wasn't because the AI was stupid or didn't understand medicine. Instead, it was because the AI was using a hidden, default rule that valued saving resources over saving lives. It was like a captain who refuses to fire the warning shot unless the asteroid is 99% certain, even though missing one would be catastrophic. The paper shows that when the researchers explicitly told the AI, "Hey, missing an emergency is 5 times worse than a false alarm," the AI instantly changed its mind. It started sending more patients to the ER, catching about 50% more real emergencies without causing a huge spike in unnecessary visits.

The study suggests that the problem isn't the AI's ability to predict the future, but its ability to follow the captain's orders on what matters most. The researchers found that by simply changing the words in the prompt to specify the "cost" of different mistakes, they could steer the AI to act like a safety-first doctor or a resource-saving doctor, depending on what was needed. They also noted that without these clear instructions, every AI model seems to have its own weird, invisible preference that no one can see or control. The paper concludes that for these tools to be safe and useful in real hospitals, we can't just ask them to "be smart." We have to explicitly tell them how to weigh the risks, just like a captain must give clear orders to their navigation computer before entering the fog.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →