Translating the Untranslatable: An Operationalizable Ontology for Untranslatability
This paper introduces a structured ontology and taxonomy of untranslatability to operationalize the concept for NLP, resulting in a new multilingual dataset and initial findings that demonstrate translation quality improves when models employ explanatory compensation strategies.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: When Translation Hits a Wall
Imagine you are trying to move a piece of furniture from your house to a new one. Usually, you just pick it up and carry it over. That's how most computer translation works today: it moves words from Language A to Language B.
But sometimes, the furniture is shaped weirdly, or the doorways in the new house are too small. You can't just carry it through; you have to take it apart, wrap it in bubble wrap, or even leave a note explaining why it looks different.
This paper is about those "weirdly shaped" moments in language. The authors call this untranslatability. It happens when a word or phrase in one language (like Japanese or Spanish) carries a meaning, a joke, or a cultural feeling that simply doesn't exist in English. If you try to translate it word-for-word, you lose the soul of the message.
The Problem: Computers Are Too Literal
Current AI translators are like very strict librarians. They are great at finding the exact dictionary definition of a word. But they struggle when a word is more like a feeling, a pun, or a cultural inside joke.
For example, the paper mentions the Japanese internet slang "wwww" (which looks like grass and means "lol"). An AI might try to translate it as "grass," which makes no sense to an English speaker. Or it might translate a Spanish word for "time spent chatting after a meal" (sobremesa) as just "after dinner," losing the cozy, social vibe of the original word.
The authors argue that we need to stop treating these as "bugs" or rare mistakes. Instead, we should treat them as a specific category of problem that needs a special toolkit.
The Solution: A New Map and a Toolkit
To fix this, the researchers built two main things:
1. The Map (The Ontology)
They created a structured map to categorize why something is untranslatable. They broke it down into three main neighborhoods:
- Linguistic: The grammar or sounds are different (like a tongue-twister that doesn't work in another language).
- Figurative: It's about jokes, puns, or idioms (like saying "it's raining cats and dogs").
- Cultural: It's about history, religion, or customs (like a greeting that involves kissing on the cheek, which might be weird in another culture).
2. The Toolkit (Compensation Strategies)
Since you can't just "carry" the meaning over, you need a strategy to get the message across. The authors defined six ways to handle this, like different tools in a toolbox:
- Annotation (The "Note" Strategy): Translate the word, but add a little explanation in parentheses. Example: "We enjoyed the sobremesa (a long chat after dinner)."
- Adaptation (The "Swap" Strategy): Replace the original joke with a different joke that makes the audience laugh in the same way, even if the words are totally different.
- Borrowing (The "Keep It" Strategy): Just keep the original foreign word and hope people learn it. Example: "We had a nice feng-shui."
- Paraphrase (The "Explain It" Strategy): Throw away the literal words and just explain the meaning in plain English.
- Options (The "Maybe" Strategy): If the original sentence is vague, give the reader multiple choices. Example: "He/She went home."
- Calque (The "Literal Copy" Strategy): Translate the words exactly as they are, even if it sounds weird, so the reader can guess the meaning.
The Experiment: Testing the Tools
The researchers didn't just talk about this; they built a massive dataset. They took 1,300 tricky sentences from Spanish and Japanese and generated 18,200 different English translations, each using a different strategy from their toolkit.
Then, they asked real humans to act as judges. They showed people the original sentence and the different translations, asking: "Which one do you like better?"
What They Found
The results were surprising and important:
- People love the "Note" strategy: The most popular translation was usually the one that included an explanation (Annotation). People preferred knowing why a word was used, even if it made the sentence longer.
- Context matters: If the translation was for a textbook (where learning is key), people loved the explanations. If it was for a movie (where flow is key), they wanted it shorter, but they still preferred explanations over confusing literal translations.
- One size does not fit all: A strategy that works for a pun might fail for a cultural custom. The best translation depends on what is being translated and where it will be read.
The Takeaway
The paper concludes that machine translation is currently trying to force a square peg into a round hole. By recognizing that some things cannot be translated directly, and by giving computers a structured list of strategies (like adding notes or swapping jokes) to handle those moments, we can make translations that actually make sense to humans.
They aren't just fixing errors; they are teaching computers how to be more like human translators: flexible, creative, and willing to add a little extra context to make sure the message is truly understood.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.