← Latest papers
💬 NLP

When Semantic Overlap Is Not Enough: Cross-Lingual Euphemism Transfer Between Turkish and English

This study investigates cross-lingual euphemism transfer between Turkish and English, revealing that semantic overlap alone is insufficient to guarantee positive transfer and that performance can sometimes improve with non-overlapping terms due to differences in label distribution and pragmatic context.

Original authors: Hasan Can Biyik, Libby Barak, Jing Peng, Anna Feldman

Published 2026-02-20
📖 5 min read🧠 Deep dive

Original authors: Hasan Can Biyik, Libby Barak, Jing Peng, Anna Feldman

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to understand human "soft talk."

In our daily lives, we often use euphemisms—polite words to replace harsh or embarrassing truths. Instead of saying "He was fired," we say "He was let go." Instead of "He died," we say "He passed away." These words are tricky because they rely heavily on culture and context, not just dictionary definitions.

This paper asks a big question: If we teach a computer to spot these polite lies in English, will it automatically know how to spot them in Turkish?

To find out, the researchers built a special experiment using two languages: English (a language the computer knows very well) and Turkish (a language the computer knows less well).

The Two Types of "Soft Talk"

The researchers realized that not all polite words are created equal. They sorted them into two buckets:

  1. The "Twins" (Overlapping PETs): These are concepts where both languages use a similar "soft" strategy.
    • Example: In English, we say someone "passed away." In Turkish, they say "vefat etmek" (which literally means "to fulfill a duty" or "to depart"). Both languages treat death like a journey or a departure. The computer can easily see the pattern here.
  2. The "Strangers" (Non-Overlapping PETs): These are concepts where one language has a polite word, but the other just uses the blunt, direct word.
    • Example: In English, if someone is unemployed, we might say they are "between jobs." In Turkish, there isn't a cute, idiomatic phrase for this; they just say "unemployed." If the computer learns the English "between jobs" trick, it won't know what to do when it sees the direct Turkish version because the "soft" version doesn't exist there.

The Big Surprise: Knowing More Doesn't Always Help

The researchers trained a smart AI model (XLM-R) to spot these words. Here is what they found, explained through a few analogies:

1. The "Overfitting" Trap (When learning too much hurts)
Imagine you are a student who memorizes the answers to a specific math test in English. If you take the same test in Turkish, you might do okay because the math is the same. But if you study too hard specifically for the English test, you might actually forget the general rules of math!

  • The Result: When the AI was trained specifically on English "soft talk," it sometimes got worse at spotting Turkish "soft talk" than if it had just been given a general, untrained brain. It learned English-specific tricks that didn't apply to Turkish.

2. The One-Way Street
The transfer of knowledge was very uneven.

  • English → Turkish: The AI was great at this. Because English has so much data, the AI learned the concept of being polite so well that it could spot it in Turkish, even when the Turkish words were totally different.
  • Turkish → English: This was a disaster. When the AI tried to learn from the smaller Turkish dataset and apply it to English, it struggled badly. It's like trying to teach a child to drive a Ferrari using only a bicycle. The Turkish data wasn't "rich" enough to teach the AI the complex rules of English politeness.

3. The "Category" Problem
The AI did okay with general topics like "Death" (because both cultures have similar polite ways to talk about it). But it completely failed on specific topics like Politics or Employment.

  • Why? In politics, English has very specific, culture-bound euphemisms (like "regime change" or "between jobs") that simply don't exist in Turkish. The AI tried to force a connection where there was none, leading to confusion.

The "Magic" vs. The "Robot"

The researchers also tested a giant, famous AI (GPT-4o) without training it on their specific data.

  • The Giant AI: It was good at guessing, but it had a weird bias. It just assumed everything was a euphemism because it was too eager to be polite. It got lucky on some tests because its "guessing bias" matched the data, not because it truly understood the language.
  • The Trained Robot: The smaller, trained model was actually more reliable. It proved that you don't need a massive, expensive brain to do this job; you just need to train a smaller brain carefully on the right kind of data.

The Takeaway

The main lesson of this paper is: Just because two languages share a meaning, it doesn't mean a computer can easily translate the "vibe" between them.

  • If you teach a computer to be polite in a language with lots of data (English), it can often figure out how to be polite in a language with less data (Turkish).
  • But if you try to teach it from the small language to the big one, it often fails.
  • And sometimes, teaching it too specifically about one language makes it forget how to generalize to others.

In short, teaching a computer to understand human "soft talk" is less like translating a dictionary and more like teaching it to understand the unspoken rules of a party. You can't just translate the words; you have to understand the culture, and that's much harder to do across different languages.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →