Do LLMs Truly Benefit from Longer Context in Automatic Post-Editing?
This paper systematically evaluates LLMs for automatic post-editing and finds that while proprietary models achieve near-human quality with simple prompting, they fail to effectively utilize document-level context, suffer from high costs, and rely on metrics that poorly reflect their actual performance improvements.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a professional translator who just received a rough draft of a novel written by a clumsy robot. Your job is Post-Editing: you read the robot's draft and fix the mistakes to make it sound natural and accurate.
For a long time, translators worked sentence-by-sentence, like fixing one sentence in a paragraph and then moving to the next without looking back. But now, we have Large Language Models (LLMs)—super-smart AI brains—that can read an entire book at once. The big question this paper asks is: "Does letting the AI read the whole book (the 'long context') actually help it fix the translation better, or is it just a waste of time and money?"
Here is the story of what they found, explained with some everyday analogies.
1. The Experiment: The "Blind" vs. The "Bookworm"
The researchers set up a race between two types of AI editors:
- The "Blind" Editor (APEseg): This AI only sees the specific sentence it needs to fix. It's like a mechanic fixing a car engine while wearing a blindfold, only looking at the one bolt in front of them.
- The "Bookworm" Editor (APEdoc): This AI sees the sentence plus the entire surrounding document. It's like a mechanic who can see the whole car, the garage, and the owner's manual all at once.
They tested this on two types of AI:
- The "Super-Models" (Proprietary): Like GPT-4o. These are expensive, closed-source giants (like a highly paid, elite consultant).
- The "Open-Source" Models: Like LLaMA or Qwen. These are free to download but smaller and less polished (like a talented but inexperienced intern).
2. The Big Surprise: "More Context" Didn't Mean "Better Results"
You might think, "If the AI reads the whole book, it will understand the story better and fix the translation perfectly!"
The Reality:
- For the Super-Models: They were already so good at fixing sentences that reading the whole book didn't really help them improve much. They were like a master chef who can make a perfect omelet with just one egg; giving them the whole fridge didn't make the omelet any tastier.
- For the Open-Source Models: Giving them the whole book actually confused them. Instead of fixing the sentence, they started hallucinating (making things up) or rewriting the whole thing incorrectly. It's like giving a nervous intern a 500-page manual; they got overwhelmed and started scribbling nonsense on the page.
The Takeaway: Just because an AI can read a long document doesn't mean it knows how to use that information to fix a single sentence.
3. The "Data Poisoning" Problem
The paper found that when you feed these models a long document, it's like throwing a bunch of noise into a quiet room.
- The Super-Models are like a noise-canceling headphone user. They ignore the background noise (the rest of the document) and focus on the task. They are robust but stubborn; they rarely use the extra context to fix subtle errors.
- The Open-Source Models are like a person with no headphones. The background noise drowns out the instructions. They get distracted by irrelevant parts of the document and start "hallucinating"—saying things that weren't in the original text at all.
4. The Cost of "Reading the Whole Book"
This is the most practical part of the paper.
- The Price Tag: Asking the AI to read the whole document is incredibly expensive. It's like asking a taxi driver to drive around the entire city to drop you off at a house just down the street.
- For the Super-Models, the cost went up by 4,300% (that's 43 times more expensive!) just to get the same result.
- For the Open-Source models, it took 10 times longer to process, and they crashed or gave bad answers more often.
5. The "Invisible" Improvements
The researchers found that sometimes the "Bookworm" editor did make the translation sound more natural (like changing a stiff phrase to a casual one). However, standard computer tests (metrics) couldn't see this improvement. It's like a human tasting a soup and saying, "This tastes better," but a machine measuring the salt content saying, "The salt is exactly the same."
This proves that humans are still needed to judge if a translation is truly good. Computers can't always tell the difference between a "good" edit and a "bad" edit when the meaning stays the same.
Summary: What Should We Do?
The paper concludes with a few key lessons:
- Don't force the AI to read the whole book yet. Currently, the "naive" approach of just dumping the whole document into the AI's prompt is too expensive and often confusing.
- The "Super-Models" are great, but expensive. They can fix translations almost as well as humans, but they cost too much to run for every single sentence.
- The "Open-Source" models need training. They are currently too easily distracted by long texts. We need better ways to teach them how to focus on the right parts of a document without getting overwhelmed.
- We need new tools. Instead of just feeding the AI the whole book, we need smarter ways to feed it only the relevant parts (like a librarian handing you just the specific chapter you need, not the whole library).
In a nutshell: Giving AI a "long context" sounds like a great idea, but right now, it's like giving a child a whole encyclopedia to help them tie their shoes. It's too much information, it costs too much to process, and it often leads to more mistakes than solutions. We need smarter ways to use that context before it becomes a practical tool.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.