An Investigation of Translationese in the Generations of Multilingual Large Language Models
This paper investigates whether multilingual large language models (MLLMs) produce "translationese" resembling human-translated text by analyzing linguistic features across five languages and comparing MLLM generations against both non-translated and human-written baselines.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
When a person speaks or writes in a language that is not their native tongue, their words often carry a subtle, invisible fingerprint. They might use sentence structures that feel slightly stiff, choose words that are technically correct but rarely used by locals, or arrange phrases in a way that betrays the influence of their first language. Linguists call this phenomenon "translationese." It is the lingering echo of the original language, visible even when the text has been successfully translated. For decades, this concept has been studied in human translation, where scholars analyze how the act of translating changes the texture of a language. Today, however, a new question has emerged in the field of artificial intelligence. Large language models, the powerful computer systems that can write fluently in dozens of languages, are becoming increasingly common. While these machines can produce text that looks perfect on the surface, researchers are beginning to wonder if they are truly thinking in the language they are writing, or if they are secretly translating everything from a dominant internal language, likely English, before outputting the final result. If they are doing this, their output would carry the same hidden fingerprints of translation that human translators leave behind.
A team of researchers set out to investigate this possibility by treating the text generated by these artificial intelligence models as if it were a translation. They wanted to know if the machines were producing text that sounded like it had been translated, even when they were asked to write directly in a specific language. To do this, they gathered text samples in five different languages: English, German, Spanish, Greek, and Pashto. They created a massive collection of writing that included three distinct types: text written naturally by native speakers, text that was genuinely translated by humans, and text generated by two of the most advanced AI models available. They also included a fourth type, where the AI wrote in English and a separate computer program translated it into the target language, to see how that compared to the AI writing directly in the target language. The researchers then used a sophisticated computer program trained to recognize the specific patterns of translationese. This program had been taught on thousands of examples of human writing and human translations, learning to spot the subtle statistical differences that separate a native speaker from a translator.
The results revealed a clear pattern. When the AI models wrote in English, their text was rarely flagged as translated, suggesting they were operating naturally in their dominant language. However, as soon as the models switched to other languages, the story changed. In German and Spanish, the AI-generated text showed strong signs of translationese. The computer program identified a significant portion of the German text and a large amount of the Spanish text as having the distinct markers of translation. This suggests that even when the models are prompted to write directly in these languages, they may be relying on an internal process that resembles translation, perhaps thinking in English first and then converting the ideas. The effect was even more pronounced when the researchers looked at the specific words the models chose. They found that the AI tended to avoid the most obvious, clumsy mistakes that humans might make, but it still struggled with the subtle distribution of common words. For instance, in German, the models used certain common connecting words in patterns that matched translated text rather than native speech, while in Spanish, they used articles like "the" far less frequently than a native speaker would.
Interestingly, the degree of this "translation" effect depended heavily on the language itself. The models performed better in Greek and Pashto, two languages with fewer digital resources available for training, than they did in German and Spanish. In these lower-resource languages, the AI-generated text was less likely to be flagged as translated. The researchers suggest this might be because the training data for German and Spanish contains a vast amount of existing translated material, which the models simply mimic. In contrast, the training data for Greek and Pashto may contain less translated content, forcing the models to rely more on the native structures available to them. The team also asked human readers to judge a small sample of the text. Surprisingly, the humans were often unable to detect the translationese that the computer program found so easily. The humans mostly noticed strange word choices or odd combinations, whereas the computer detected the deeper statistical shifts in how words were arranged. This indicates that while the AI might be hiding the most obvious errors, the underlying structure of its writing still betrays a lack of native fluency.
The study concludes that these powerful language models do not yet generate text in a truly native way for non-English languages. Instead, their output often carries the hidden signature of translation, suggesting that their internal processing may still be anchored to English. This does not mean the models are failing; they are often highly effective and useful. However, it reveals a fundamental limitation in how they learn and operate. They appear to be mastering the surface level of many languages while still relying on a central, English-based framework for their reasoning. The researchers note that this finding helps explain why these models sometimes struggle with cultural nuances or the natural flow of a language that is not English. By identifying these hidden fingerprints, the study provides a new way to measure the true quality of machine-generated text, moving beyond simple correctness to assess the genuine naturalness of the voice the machine is trying to emulate.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.