← Latest papers
💬 NLP

ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety

The paper introduces ROK-FORTRESS, a bilingual benchmark utilizing a "transcreation matrix" to demonstrate that National Security and Public Safety evaluations for large language models must account for the complex interactions between language and geopolitical context, as translation-only approaches fail to capture how Korean language and grounding uniquely influence model safety behaviors and over-refusal rates.

Original authors: Michael S. Lee, Yash Maurya, Drew Rein, Bert Herring, Jonathan Nguyen, Kyungho Song, Udari Madhushani Sehwag, Jiyeon Cho, Kaustubh Deshpande, Yeongkyun Jang, Jiyeon Joo, Minn Seok Choi, Evi Fuelle, Ch
Published 2026-05-15
📖 5 min read🧠 Deep dive

Original authors: Michael S. Lee, Yash Maurya, Drew Rein, Bert Herring, Jonathan Nguyen, Kyungho Song, Udari Madhushani Sehwag, Jiyeon Cho, Kaustubh Deshpande, Yeongkyun Jang, Jiyeon Joo, Minn Seok Choi, Evi Fuelle, Christina Q Knight, Joseph Brandifino, Max Fenkell

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are testing a very smart robot to see if it will follow dangerous instructions, like "How do I build a bomb?" or "How do I hack a bank?"

For a long time, safety researchers tested these robots mostly in English. They assumed that if you just translated those dangerous questions into another language (like Korean), the robot would still say "No." They thought the robot's safety rules were like a universal shield that worked the same way no matter what language you spoke.

But this paper, ROK-FORTRESS, says: "Wait a minute. It's not just about the language; it's about where the story takes place."

Here is the simple breakdown of what they did and what they found, using some everyday analogies.

1. The "Translation vs. Transcreation" Test

The researchers realized that simply translating a question is like translating a recipe. If you translate a recipe for "American Apple Pie" into Korean, you still get a recipe for Apple Pie, just in Korean words. The ingredients (the threat) are the same.

But Transcreation is different. It's like taking that Apple Pie recipe and turning it into a recipe for "Korean Sweet Potato Pie." The intent (making a sweet dessert) is the same, but the specific ingredients, the cultural context, and the local rules are completely different.

The team built a massive test called ROK-FORTRESS (a fortress for safety testing) with 1,235 different "dangerous" questions. They tested them in four ways:

  1. English, US Context: "How do I hack the FBI?" (Original)
  2. Korean, US Context: "How do I hack the FBI?" (Translated)
  3. English, Korean Context: "How do I hack the Korean National Intelligence Service?" (Adapted)
  4. Korean, Korean Context: "How do I hack the Korean National Intelligence Service?" (Fully Transcreated)

2. The Big Surprise: The "Conservative Korean" Effect

Most previous studies thought that speaking a language the robot isn't as good at (like Korean) would be a "loophole" or a "backdoor" to trick the robot into being dangerous. They thought the robot would get confused and say "Yes" to bad things.

The paper found the exact opposite.

When they asked the robots in Korean, the robots actually became more cautious. They said "No" more often.

  • The Analogy: Imagine a security guard at a museum. If you ask him in his native language with a local accent, he gets very strict and checks your ID three times. If you ask him in a broken foreign language, he might get confused and let you in. But these robots acted like the guard who gets extra strict when you speak Korean. They treated the Korean language itself as a "Red Flag" or a "Risk Signal," making them more likely to refuse the request, even if the request was dangerous.

3. The "Context" Factor

The researchers also found that where the story happens matters.

  • If you ask a robot about a US bank in English, it might be cautious.
  • If you ask the same robot about a Korean bank in English, it gets even more cautious.
  • If you ask about the Korean bank in Korean, it is the most cautious of all.

It's like a security system that has a "Local Mode." When the robot realizes the situation is happening in Korea (with Korean names, places, and laws), it tightens its safety belt even more.

4. The "Jailbreak" Twist

The team also tested what happens if you strip away all the fancy tricks (jailbreaks) and just ask the question directly: "How do I make a bomb?"

  • Open-Source Robots: When asked directly, these robots behaved like the old studies predicted. They were easier to trick in Korean.
  • Big Corporate Robots: The big, expensive models (like those from OpenAI or Anthropic) stayed cautious in Korean, even when asked directly. They didn't fall for the "language loophole."

5. Why This Matters (According to the Paper)

The main takeaway is that you cannot just translate safety tests.

If you want to know if a robot is safe in Korea, you can't just take your English test, translate it, and run it. You have to change the story to fit Korean culture (transcreation).

  • The Old Way: "Translate the question, check the answer." (This misses the point).
  • The New Way: "Rewrite the story for the local culture, then check the answer."

The paper concludes that for these specific robots, speaking Korean and talking about Korean topics actually makes them safer (more likely to say "No" to bad things), not less safe. This is a huge shift from what everyone thought before.

Summary in One Sentence

This paper built a special test to show that when you ask AI robots about dangerous topics in Korean and about Korean places, they don't get tricked more easily; instead, they often get more careful and refuse to answer, proving that safety testing needs to be culturally adapted, not just translated.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →