← Latest papers
💬 NLP

When Readability and Source Retention Diverge: An Evaluability Gap in AI Translation

This study reveals that while readable AI translations often create an "evaluability gap" where users fail to recognize fidelity differences even when the source text is visible, their trust and willingness to disclose personal information are more strongly driven by perceived task performance and anthropomorphic attribution than by objective content retention.

Original authors: Chenchen Mao, Hanjing Shi, Haiyan Jia, Emily Wegrzyn, Dominic DiFranzo

Published 2026-08-20
📖 5 min read🧠 Deep dive

Original authors: Chenchen Mao, Hanjing Shi, Haiyan Jia, Emily Wegrzyn, Dominic DiFranzo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the quiet corners of our digital lives, we constantly hand over pieces of ourselves to machines. We ask translation tools to carry our private messages across language barriers, trusting them to keep our meaning intact while making it understandable to someone else. For years, the goal of these tools was simply to produce text that flowed smoothly, sounding natural and easy to read. But as artificial intelligence has become more powerful, a new tension has emerged. A machine can now rewrite a difficult text to make it simpler and more readable, but in doing so, it might accidentally strip away the original meaning or leave out crucial details. This creates a tricky situation for the person using the tool: they can see the original text on the screen, yet they might not be able to tell if the machine's version has changed the story. The question researchers are now asking is whether simply showing the original text is enough to help us judge if a translation is good, or if there is a gap between what we can see and what we can actually understand.

A team of researchers at Lehigh University set out to explore this gap by testing how people judge translations when the machine prioritizes either easy reading or strict accuracy. They built a simple, text-only interface called TransLingo, designed to look like a standard utility without any friendly avatars or voices that might trick a user into thinking the machine is human. They invited over three hundred people to use this tool with two different types of source text. The first type was simple, everyday narrative prose, similar to a story a child might read. The second type was complex, literary-philosophical writing, drawn from ancient texts that require a college-level education to fully grasp. For each type of text, the researchers presented two versions of the translation. One version was generated by an artificial intelligence to be as readable and smooth as possible, even if it meant simplifying the content. The other version was carefully revised by a human researcher to stay as faithful as possible to the original words and ideas, even if it remained slightly more difficult to read.

The researchers wanted to see if people could spot the difference between these two versions and if their judgment changed depending on how hard the original text was to understand. When the source text was simple, the participants were sharp. They could clearly see that the version which stayed true to the original content was better. They rated the faithful translation higher in quality, and they also felt that the machine itself was smarter and more trustworthy when it produced that accurate version. However, the results took a surprising turn when the source text was complex. Even though the researchers had verified that the faithful version still contained more of the original meaning, the participants could not tell the difference. They rated the easy-to-read version and the faithful version as equally good. In fact, for the difficult texts, the machine's ability to make the output readable seemed to blind the users to the fact that the faithful version was preserving more of the source material. The source text was right there on the screen, visible for comparison, yet the users' overall judgment did not reflect the difference in what was actually retained.

This phenomenon, which the authors call an "evaluability gap," suggests that showing a user the original text is not enough to guarantee they can evaluate the translation correctly. When the material is complex, the effort required to compare the two texts seems to push users toward relying on how easy the output feels to read, rather than checking if the meaning has been preserved. The study also looked at how these judgments affected trust and the willingness to share personal information. The researchers found that when people trusted the machine to perform its task accurately, they were more willing to say they would share personal details like their political views, financial status, or family relationships with it. However, the type of translation they saw did not change their willingness to share. Whether the machine produced a readable or a faithful translation, the participants' stated comfort with disclosing personal information remained the same. This indicates that judging the quality of a translation and deciding whether to entrust a machine with private data are two separate mental processes.

The findings challenge the idea that a transparent interface, where the original text is always visible, solves the problem of trust in artificial intelligence. The study shows that for complex tasks, visibility does not equal understanding. Users may look at the source and the output side by side, but if the output is made to feel smooth and easy, they may overlook the fact that the machine has altered the content. This has important implications for how we design tools that handle our personal information. It suggests that we cannot rely on users to figure out if a machine is being honest just by giving them the source text. Instead, designers may need to provide specific help, such as highlighting exactly what was changed or simplified, to bridge the gap between seeing the text and truly understanding what the machine has done. The research confirms that while we can build machines that read well, ensuring that we can also judge what they have kept is a much harder problem to solve.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →