← Latest papers
💬 NLP

Fluency and Faithfulness in Human and Machine Literary Translation

This study analyzes over 130,000 translated paragraphs from 106 novels to reveal a consistent negative correlation between fluency and faithfulness in literary translation for both humans and Google Translate, while noting that this trade-off is weaker or non-significant for TranslateGemma.

Original authors: Sarah Griebel, Ted Underwood

Published 2026-05-18
📖 5 min read🧠 Deep dive

Original authors: Sarah Griebel, Ted Underwood

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to translate a complex, beautiful novel from a foreign language into English. You have two main goals, but they often pull you in opposite directions:

  1. Fluency: Making the English sound smooth, natural, and like it was written by a native speaker.
  2. Faithfulness: Keeping the meaning and structure exactly as the original author intended, even if it sounds a little "foreign" or awkward.

This paper asks a simple question: In the world of literary translation, do these two goals fight each other? Does making a translation sound more natural automatically mean you are losing some of the original meaning?

The authors, Sarah Griebel and Ted Underwood, decided to test this using a massive library of 130,000 paragraphs from 106 different novels. They compared three types of translators:

  • Humans: Professional human translators.
  • Google Translate: A traditional machine translation system.
  • TranslateGemma: A newer, advanced AI (Large Language Model).

The Tools: How They Measured the "Tug-of-War"

To measure Fluency, they didn't ask humans to read the text. Instead, they built a "detective bot." This bot was trained to spot the difference between text written originally in English and text that had been translated.

  • The Trick: To make sure the bot wasn't just guessing based on what the story was about (the vocabulary), they stripped the sentences down to their skeleton: the parts of speech (nouns, verbs, adjectives, etc.).
  • The Logic: If a paragraph looks like a native English skeleton, the bot gives it a high "Fluency" score. If it looks like a translated skeleton (a bit clunky or structured differently), the score is lower.

To measure Faithfulness, they used a tool called COMET-KIWI. Think of this as a "meaning checker." It looks at the source text and the translation and tries to guess: "How much of the original meaning did this keep?" It doesn't need a perfect human translation to compare against; it just evaluates the pair directly.

The Big Discovery: The "Length" Trap

Before finding the main answer, the researchers hit a snag. They noticed that paragraph length was messing up their results.

  • The Analogy: Imagine trying to judge how well a student summarized a book. If the student writes a summary that is only one sentence long, it's hard to be faithful to the whole story, but it's easy to make that one sentence sound smooth. If they write a 10-page summary, it's harder to keep it perfectly smooth, but easier to keep all the details.
  • The Finding: Long paragraphs tended to get lower scores on both fluency and faithfulness. Short paragraphs got higher scores on both. This made it look like fluency and faithfulness were unrelated, when really, the length of the text was the hidden variable pulling the strings.

The Real Answer: The Trade-Off

Once they mathematically "controlled" for the length of the paragraphs (comparing apples to apples), a clear pattern emerged: There is a trade-off.

When a translation becomes more fluent (more native-sounding), it tends to become slightly less faithful to the original meaning.

  • Human Translators: They showed the strongest version of this trade-off. When humans made the text sound very natural, they often had to tweak the meaning slightly to do it.
  • Google Translate: Also showed this negative relationship, though slightly weaker.
  • The AI (TranslateGemma): This was the surprise. The AI showed a much weaker, sometimes even non-existent trade-off. It seemed to be able to be fluent and faithful at the same time, or at least, the relationship between the two wasn't as clear-cut as it was for humans.

What Does This Mean?

The paper suggests that in literary translation, you can't always have it all.

  • If you want the text to read perfectly like a native English novel, you might have to sacrifice a tiny bit of the original structure or meaning.
  • If you want to stick rigidly to the original structure to preserve every nuance, the text might sound a bit "foreign" or less smooth.

The authors also note that their "Fluency" detector was trained on books from the 1800s and early 1900s. This means it favors the style of older English literature. Since human translators often try to match that historical style, they might score higher on fluency. Modern AI, which is trained on modern internet text, might sound "too modern" to the detector, even if it's grammatically perfect.

Summary

In short, the paper proves that in the world of translating novels, smoothness and strict accuracy are often in a gentle tug-of-war. Making a translation sound more like a native English book often requires a small compromise on how strictly it follows the original. Interestingly, while humans and older AI systems show this tension clearly, newer, smarter AI models seem to handle this balancing act differently, sometimes managing to be both smooth and accurate without the same clear trade-off.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →