← Latest papers
💬 NLP

LLM-Powered Automatic Translation and Urgency in Crisis Scenarios

This paper evaluates the performance of large language models and machine translation systems in crisis scenarios, revealing significant quality degradation in low-resource languages and a critical inconsistency in LLM-based urgency classification across different languages that poses risks for effective crisis triage.

Original authors: Belu Ticona, Antonis Anastasopoulos

Published 2026-08-13
📖 5 min read🧠 Deep dive

Original authors: Belu Ticona, Antonis Anastasopoulos

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where a giant, super-smart robot librarian can read every book in every language instantly. This is the promise of Large Language Models (LLMs), a type of artificial intelligence that has learned to speak, write, and reason by reading billions of sentences. Think of them as digital polyglots who can summarize a news article, write a poem, or translate a recipe from French to Japanese in a split second. But here's the catch: most of these robots were trained mostly on English books and websites. They are like a chef who is a master at making French cuisine but has never tasted a single dish from the rest of the world.

Now, picture a crisis. A flood is rising, or an earthquake has just hit. In these moments, every second counts, and the difference between a "get out now" message and a "stay put" message can mean life or death. Emergency responders need to translate urgent warnings into local languages instantly to save lives. They hope to use these super-smart robot librarians to do the heavy lifting. But the big question is: if you ask a robot trained mostly on English to translate a life-or-death warning into a rare local language, will it get the meaning right? Will it understand that "urgent" means "run now," or will it accidentally turn it into "maybe later"? This is the dangerous gap this paper explores.

The researchers, Belu Ticona and Antonios Anastasopoulos, decided to put these AI translators to the ultimate test. They didn't just ask, "Can it translate?" They asked, "Can it translate urgency correctly?" Crucially, they focused their study on open-weight models—the versions of AI that are freely available for anyone to download, inspect, and run on their own computers. This is a key choice because these are the tools actually accessible to humanitarian NGOs, local governments, and crisis responders in underserved regions who can't afford expensive commercial subscriptions or rely on unstable internet connections. They treated these specific, open-source AI recruits like new hires in a disaster zone and gave them a massive pile of crisis-related texts—like news reports and safety instructions—covering 30 different languages from all over the globe, including many that are rarely seen on the internet.

First, they looked at the translation quality. Imagine trying to translate a complex instruction manual into a language the robot barely knows. The results were a bit of a rollercoaster. For languages the robots knew well, the translations were decent. But for many languages from Africa and parts of Asia, the robots completely stumbled. In some cases, the translation scores dropped so low that the meaning was almost totally lost. It was as if the robot tried to translate a warning about a rising flood and ended up saying something about a gentle rain. The study found a particularly dangerous trend: the robots were much better at reading local languages and translating them into English than they were at taking English instructions and translating them out into local languages. This is a critical risk because it means the AI might be good at understanding what a local community is saying, but terrible at sending life-saving instructions back to them. While some models were consistent, others were wildly unpredictable; a model might do a great job for one language and a terrible job for a neighbor, even if they are in the same "multilingual" family.

Then, the researchers dug deeper into the "urgency" problem. They created a special set of 100 crisis scenarios, ranging from "not urgent" to "critical," and translated them into 29 languages. They asked both human experts and the AI to read these translated warnings and decide how urgent they were. Here is where things got really strange. The humans were rock-solid; no matter what language they read, they all agreed on how dangerous the situation was. If a human read a warning in Spanish, Greek, or Hindi, they all said, "This is Critical!"

But the AI? The AI was all over the place. For the exact same warning, the AI might call it "Critical" in Spanish, "Medium" in Bengali, "Very Low" in Greek, and "Not Urgent" in Hindi. It was like a mood ring that changed color depending on the language it was wearing. The AI didn't just make small mistakes; it swung wildly across the entire scale of urgency. A message that humans agreed was a life-or-death emergency was sometimes labeled by the AI as something you could ignore.

The paper suggests that this isn't just a minor glitch; it's a serious risk. The researchers found that the AI's ability to judge urgency depends heavily on the language it's reading, and for many low-resource languages, it simply cannot be trusted to get the tone right. They argue that we can't just plug these general-purpose AI tools into emergency response systems without checking them first. If a robot thinks a critical evacuation order is "not urgent" just because it's written in a specific local dialect, people could get hurt.

In short, the study shows that while these AI models are impressive, they are currently too inconsistent to be the sole heroes of a crisis. They might work well for some languages, but for others, they might fail to understand the difference between a "heads up" and a "run for your life." The authors conclude that before we let AI handle crisis communication, we need to test it rigorously in the specific languages and situations where it will be used, and we must keep human experts in the loop to make sure the urgency isn't lost in translation.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →