LLMs in the Real World: Evaluating "AI" in Emergency Contexts
This paper calls for researchers to better communicate their findings to the public by using a case study of an LLM-based text-to-911 translation system to highlight common misconceptions about AI in emergency contexts and to provide concrete recommendations for stakeholders throughout the development and deployment pipeline.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: A Call to Wake Up
Imagine the world of Artificial Intelligence (AI) research as a group of brilliant architects who design incredibly complex, futuristic bridges. These architects know exactly how the steel is made, where the weak spots are, and what happens if a hurricane hits.
However, the paper argues that these architects are staying in their towers, while the people who actually build and drive on these bridges (city officials, police, emergency services) are being sold a brochure that says, "This bridge is magic! It works for everyone!"
The authors, Sara Court, Lara Downing, and Micha Elsner, are urging the researchers to come down from the tower and talk to the public. They want to stop the "hype" and explain the real risks before someone gets hurt.
The Case Study: The "Magic" 911 Translator
To prove their point, the authors looked at a real-life situation in a US city. The city launched a new Text-to-911 service.
- The Promise: The city advertised that you could text 911 in 55 different languages. The idea was that if you couldn't speak English, you could just text in your native tongue, and a computer would instantly translate it for the dispatcher.
- The Reality: The researchers went to the 911 center to test it. They found that the system was like a magic 8-ball that only works on sunny days.
Here is what went wrong:
- The "One Size Fits All" Trap: The system claimed to support "Arabic." But Arabic has many dialects. The system only understood Modern Standard Arabic (the formal language used in news and books). If a person texted in a local dialect or used slang, the computer got confused. It's like buying a dictionary that only has words from the 1800s and trying to order a pizza with it.
- The Script Problem: The system claimed to support "Nepali." But the version of Nepali the system understood was written in a specific script (Devanagari). However, many refugees in that area write Nepali using the Latin alphabet (like English letters) because they don't have the special keyboard on their phones. The system couldn't read their texts at all.
- The "Black Box" Mystery: The 911 center staff didn't know how the software worked. They didn't have a manual, they didn't know the error rates, and they didn't know which languages were actually safe to use. They were driving a car with no dashboard and no brakes, hoping the engine would hold up.
The Five Big Myths (Misconceptions)
The paper identifies five common lies people tell themselves about AI:
- "AI" is a Magic Word: People think "AI" is one single thing. The authors say it's more like calling all vehicles "cars." A bicycle, a semi-truck, and a rocket ship are all vehicles, but they work very differently. You can't use a bicycle to haul a house.
- Superhuman Brains: We think AI is smarter than humans. In reality, it's more like a parrot that has read the entire internet. It can repeat facts, but it doesn't understand them. If you ask it a tricky question in an emergency, it might confidently say the wrong thing.
- Language is Easy: People think translation is just swapping word-for-word. But language is like cooking. You can't just swap "salt" for "sugar" and expect the cake to taste the same. Context, tone, and culture matter. A machine often misses the "flavor" of the message.
- Numbers Don't Lie: Companies show off high scores (like 95% accuracy) to sell their products. But the authors say these scores are like a test taken in a quiet classroom. In the real world, with noise, stress, and typos, the score drops. A high score doesn't mean the system won't fail when it matters most.
- Technology Fixes Everything: This is called "Solutionism." It's the belief that if you have a problem, you just need a high-tech app to fix it. The authors argue that sometimes, the best solution is a human being. Replacing a professional human interpreter with a cheap, glitchy robot to save money is like replacing a surgeon with a robot vacuum to save on hospital bills.
The Missing Link: The "Accountability Gap"
The paper describes a broken chain of trust.
- The Researchers know the risks but aren't talking to the public.
- The Sellers are selling the tech without explaining the risks.
- The Buyers (like the 911 center) are buying it because they want to help, but they don't have the expertise to check if it's safe.
It's like a lemonade stand where the owner sells "100% pure juice," but the buyer doesn't know the owner is mixing in tap water and sugar. The buyer trusts the sign, but they get sick.
The Solution: What Should We Do?
The authors offer a set of "Best Practices" to fix this:
- Treat AI like Medicine: Just as you wouldn't take a new drug without a label listing side effects, you shouldn't use AI for emergencies without a "Model Card." This card should clearly say: "This works well for Spanish, but fails for Arabic dialects. Do not use for life-or-death situations without a human backup."
- Get Human Backups: If you use a robot translator, you must have a human ready to step in if the robot gets confused.
- Be Honest: Tell the person texting 911, "Hey, a computer is translating this. If it looks wrong, please call us."
- Researchers Need to Speak Up: The scientists who built the tech need to stop hiding in the lab and start explaining the dangers to the people who are buying the tech.
The Bottom Line
The paper concludes that when lives are on the line (like in a 911 call), we cannot rely on "good enough" technology. If a translation is wrong, someone could get hurt or even die.
The authors aren't saying "Don't use AI." They are saying, "Don't use AI blindly." We need to slow down, check the brakes, and make sure the technology is actually safe before we let it drive our emergency services. Until we do that, we are gambling with people's lives.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.