← Latest papers
💬 NLP

Which Tokens Need Context? A Reference-Based Analysis of Translation Responsibility Using Fertility and Entropy

This paper proposes a model-agnostic, post-hoc framework using word alignment-derived fertility and entropy to quantify how human translation selectively redistributes generative responsibility from source to context tokens, revealing that context primarily resolves ambiguity for function words without altering the overall information load of content words.

Original authors: Ramakrishna Appicharla, Baban Gain, Santanu Pal, Asif Ekbal

Published 2026-06-30
📖 5 min read🧠 Deep dive

Original authors: Ramakrishna Appicharla, Baban Gain, Santanu Pal, Asif Ekbal

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are translating a story from one language to another. When a human translator does this, they don't treat every single word the same way. Some words are like anchors; they stand on their own and don't need help from the rest of the story (like a specific name, "Paris"). Other words are like chameleons; their meaning changes completely depending on what was said just before or what comes next (like the word "it" or "he").

This paper asks a simple question: Which words actually need the "surrounding story" to make sense, and which ones don't?

The researchers wanted to find out if computer translation systems are using the surrounding story (context) the same way humans do. Instead of looking inside the computer's "brain" (which is complicated and varies by model), they built a new tool to look at the results of human translations.

Here is how they did it, using two simple concepts:

1. The Two Tools: "Fertility" and "Entropy"

To understand how words work, the authors used two metaphors:

  • Fertility (The "Workload"): Imagine every word in the original sentence has a job to do: it has to "give birth" to words in the new language.

    • A word with high fertility is a hard worker that turns into many words in the new language.
    • A word with low fertility might not turn into anything at all, or just one word.
    • The Paper's Finding: When you add the previous sentence to the story, the original sentence's words suddenly have less work to do. They "offload" some of their job to the new context.
  • Entropy (The "Predictability"): Imagine how consistent a word is.

    • Low Entropy means a word is predictable. It always does the same job, no matter the story. (Example: The name "Google" is always "Google").
    • High Entropy means a word is unpredictable. Its job changes wildly depending on the story. (Example: The word "bank" could be a river bank or a money bank).
    • The Paper's Finding: Words that rely on context (like pronouns) become more unpredictable (higher entropy) when you look at them in isolation, but the context helps settle them down.

2. The Big Discovery: "Responsibility Redistribution"

The most important finding of this paper is that adding context does not create new work; it just shifts the workload.

Think of it like a relay race.

  • Without Context: The runner holding the baton (the source word) has to run the whole distance alone.
  • With Context: The runner passes part of the distance to a teammate (the context word). The total distance of the race doesn't change, but the responsibility is shared.

The paper found that:

  • Function Words (The Chameleons): Words like pronouns ("he," "she"), helpers ("is," "are"), and connectors ("and," "but") are the ones that hand off the most work to the context. They rely heavily on the surrounding sentences to know what they mean.
  • Content Words (The Anchors): Words like nouns ("cat," "car") and proper names ("London") keep doing their own work. They rarely change their job, even when you add more sentences.

3. The "Random" Surprise

The researchers tested this with three types of context:

  1. The sentence that came before.
  2. The sentence that comes after.
  3. A completely random sentence from a different story.

The Result: Even when they gave the translator a random, unrelated sentence, the "workload" still shifted slightly. The source words still offloaded some responsibility to the random words.

  • Why? The computer (and the alignment tool they used) found tiny, accidental connections between words, even if they didn't make logical sense. This suggests that simply having extra text changes how the translation is built, even if that text isn't perfectly relevant.

4. Why This Matters

The authors built a "diagnostic baseline." Think of this as a gold standard map of how humans actually use context.

  • Before this paper: We didn't have a clear way to tell if a computer was using context correctly or just guessing.
  • Now: We have a map. If a computer translation system shifts the workload exactly like the human map (offloading pronouns to the context, but keeping nouns steady), it is working like a human. If it tries to make every word work harder or ignores the context entirely, we know it's not behaving naturally.

Summary in a Nutshell

This paper didn't build a new translator. Instead, it built a magnifying glass to watch how human translators use the surrounding story. They discovered that humans are very selective: they only call in the "context team" to help with tricky, ambiguous words (like pronouns), while leaving the solid, clear words (like names) to do their own thing. They proved that context doesn't add new information; it just helps redistribute the effort of translating the sentence.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →