← Latest papers
💬 NLP

Evaluative Judgement in Teaching AI-based Translation: A Class-room Case Study of AI-Mediated Translation and Post-Editing

This paper analyzes a classroom case study of 23 translation students to demonstrate how structured comparison of AI and machine translation systems fosters evaluative judgement, revealing that students prioritize factors like adequacy, fluency, and post-editing effort over automatic metrics when selecting outputs for post-editing.

Original authors: Gokhan Dogru

Published 2026-06-16
📖 4 min read☕ Coffee break read

Original authors: Gokhan Dogru

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a translation class as a cooking competition. Instead of asking students to just cook a meal from scratch, the teacher gives them a special challenge: taste-test four different pre-made dishes, pick the best one to finish, and explain exactly why they chose it.

This paper is a report on how 23 students in a university translation course handled that challenge when the "pre-made dishes" were created by Artificial Intelligence (AI).

Here is the breakdown of what happened, using simple analogies:

The Setup: The "Taste-Test" Assignment

The students were given a short, specialized text (like a Wikipedia article about animals or technology). They had to do five things:

  1. Cook from scratch: Translate the text themselves first (to understand the recipe).
  2. Order takeout: Ask four different AI "chefs" (two advanced chatbots and two standard translation engines) to translate the same text.
  3. Get a robot score: Run the four AI translations through a computer program that gives them a grade based on how closely they match the original words (like a spelling checker on steroids).
  4. Taste and judge: Read the translations carefully to see if they actually make sense, sound natural, and use the right technical words.
  5. Pick a winner: Choose one of the four AI versions to "fix up" (post-edit) into a final, publishable article. Crucially, they had to write a report explaining why they picked that specific one.

The Big Surprise: The Robot Score Isn't the Final Judge

The most interesting finding is that the students did not blindly trust the robot scores.

Think of the automatic score like a "popularity contest" or a "nutrition label." It tells you how many ingredients match, but it doesn't tell you if the food tastes good or if it's safe to eat.

  • The Result: In about half the cases, the AI that got the highest robot score was not the one the students picked to fix.
  • Why? The students realized that an AI could get a high score just by rearranging words in a way that looked correct to a computer but sounded weird to a human, or by missing a crucial technical term.

What Did the Students Look For?

Instead of just looking at the score, the students acted like quality control inspectors. They looked for:

  • The "Flavor" (Naturalness): Does it sound like a real person wrote it, or does it sound like a robot trying to sound human?
  • The "Ingredients" (Terminology): Did it get the scientific names of animals or tech terms right?
  • The "Workload" (Post-editing effort): If they picked this version, how much work would it be to fix it? Sometimes a version looked perfect but had hidden traps that would take hours to fix.
  • The "Structure" (Grammar): Did it keep the sentences in a logical order, or did it scramble them?

The Main Lesson: Students Became "Arbiters of Quality"

The paper argues that the goal of this class wasn't to find out which AI is the "best" in the world. The goal was to teach students how to be smart judges.

Before this, students might have thought, "The computer says this is 95% perfect, so I'll just use it."
After this assignment, they learned to say, "The computer says this is 95% perfect, but it got the medical term wrong and the sentence structure is awkward. I'm going to pick the one that scored 85% because it's easier to fix and sounds more natural."

The Takeaway

The study shows that when you teach students to compare different AI tools and force them to justify their choices, they stop being passive users who just accept whatever the machine gives them. Instead, they become critical editors who know how to spot the difference between a translation that looks good on a computer screen and one that is actually good for a human reader.

In short: The paper proves that with the right kind of homework, students learn to keep their human brains in the driver's seat, using AI as a helpful tool rather than a boss.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →