← Latest papers
💻 computer science

Beyond Mean Scores: Individual- and Trait-Level Agreement Between Human and Generative AI Raters in EFL Writing Assessment

This study reveals that while current generative AI models demonstrate high internal stability, they exhibit lower agreement with human raters, systematic severity biases, and failure to apply specific rubric rules compared to human double-marking, suggesting they are best suited for supervised, discrepancy-triggered assistance rather than autonomous replacement in consequential EFL writing assessment.

Original authors: Savaş Okyay

Published 2026-08-28
📖 4 min read☕ Coffee break read

Original authors: Savaş Okyay

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the high-stakes world of language learning, a single essay can determine whether a student advances to university or remains stuck in a preparatory class. For decades, educators have relied on human teachers to read these essays and assign grades based on complex rubrics that measure everything from grammar to the logical flow of an argument. Recently, a new contender has entered the arena: artificial intelligence. Large language models, the same powerful computer programs that can write stories or answer questions, are increasingly being proposed as automated graders. The promise is seductive: machines are fast, tireless, and consistent. But in the realm of education, speed is not enough. A grading system must do more than produce a plausible average score; it must agree with human experts on the specific details of every single student's work, apply the rules fairly, and recognize when a student has written something that simply does not belong in the exam. The question is no longer whether a machine can generate a number, but whether that number can be trusted to make life-altering decisions for learners.

A researcher set out to test this trust in a real-world setting, moving beyond simple comparisons of average scores to see how well machines and humans actually agree on individual papers. They gathered seventy-two handwritten essays from students in Turkey who were taking a proficiency exam to determine their readiness for degree studies. These students had written opinion pieces on topics ranging from the value of hard work to the balance between friends and possessions. The researcher then subjected these papers to a rigorous, end-to-end evaluation. First, six experienced human teachers graded every essay twice, using a five-part rubric that scored task achievement, organization, vocabulary, grammar, and coherence. Then, three different artificial intelligence models—specifically versions of Claude, GPT, and Gemini—read the scanned images of the same handwritten essays and graded them using the exact same instructions. Crucially, the machines received no special training on these specific essays; they had to apply the rules on the spot, just as a new human teacher would.

The results revealed a landscape far more complicated than the hope of a perfect, automated replacement for human judgment. While the artificial intelligence models were remarkably consistent with themselves—meaning if you asked the same model to grade the same essay twice, it would almost always give the exact same score—they struggled to align with the human teachers. The agreement between the two human raters was strong, but when the machines were compared to the human average, the connection was significantly weaker. Every single model tended to be harsher than the human panel, awarding lower scores across the board. This severity was not a fixed amount; the machines penalized stronger essays more heavily than weaker ones, meaning a simple mathematical adjustment could not fix the gap. One model, in particular, was nearly perfect at repeating its own decisions but was the most severe of all, demonstrating that consistency does not automatically mean fairness or accuracy.

The study also uncovered how these models handled the specific rules of the exam, particularly regarding essays that were completely off-topic. The rubric stated clearly that if a student wrote about the wrong subject, they should receive a score of zero on every dimension. The human teachers followed this rule without fail. The machines, however, did not. Even when presented with essays that had nothing to do with the prompt, the models still assigned them partial scores, failing to recognize that the content was ineligible for grading. This was a critical operational failure, showing that without human oversight, the machines could not distinguish between a poor answer and an irrelevant one. Furthermore, when the researcher looked at the specific traits of writing, such as vocabulary or grammar, the models often failed to see the differences between these categories, treating them as a blur rather than distinct skills.

Ultimately, the research suggests that these artificial intelligence tools are not ready to replace human raters in high-stakes exams where a student's future is on the line. The machines are not broken, but they are not yet reliable enough to stand alone. They function best not as independent judges, but as assistants that flag discrepancies or offer a second opinion for a human to review. The study concludes that for AI to be useful in education, it must be part of a supervised workflow where humans remain in control, checking the machine's work, enforcing the rules, and making the final call. The technology offers a powerful way to extend the reach of educators, but it cannot yet replace the nuanced judgment required to ensure that every student is graded fairly and accurately.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →