← Latest papers
📄 social_science

Beyond Automated Scoring: An Evidence-Informed Human–GenAI Complementarity Framework for Formative Writing Assessment

This study employs a mixed-methods design to demonstrate that while Generative AI and human teachers exhibit limited rating agreement, they achieve practical equivalence in most scoring dimensions and offer distinct feedback characteristics, leading to the development of a framework that positions GenAI as a complementary support for rather than a replacement of teacher judgment in formative writing assessment.

Original authors: Herford Rei Biscayno Guibangguibang, Kian C. Hesula

Published 2026-08-25
📖 4 min read☕ Coffee break read

Original authors: Herford Rei Biscayno Guibangguibang, Kian C. Hesula

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In classrooms around the world, the act of teaching writing is a deeply human endeavor. It relies on a teacher reading a student's work, understanding the struggle behind the words, and offering guidance that helps the writer grow. This process, known as formative assessment, is not about assigning a final grade but about providing feedback that shapes the next draft. For decades, educators have sought ways to make this feedback more timely and consistent, especially as class sizes have grown. Recently, a new tool has entered the conversation: generative artificial intelligence. These are computer systems capable of reading text and generating human-like responses, including critiques and suggestions. The central question for educators has not been whether these machines can read, but whether they can teach. Can a computer provide the same kind of useful, learning-focused feedback that a skilled teacher does, or is there a fundamental difference in how they see a student's work?

A team of researchers set out to answer this by putting human teachers and three different artificial intelligence systems to the test. They gathered one hundred essays written by university students learning English as a second language. These essays covered a range of skill levels and were then evaluated independently by three experienced teachers and three different AI systems. The goal was not simply to see if the machines could match the teachers' scores, but to understand how the feedback they produced compared in substance and style. The researchers used a standard scoring guide that looked at four specific areas: the content of the writing, how the ideas were organized, the use of language, and the mechanical details like spelling and punctuation.

The results revealed a surprising complexity. When the three teachers graded the essays among themselves, they agreed closely with one another, showing a high level of consistency. The three AI systems also agreed with each other almost as well. However, when the researchers compared the teachers directly to the AI systems, the agreement dropped significantly. The scores given by the humans and the machines often did not line up, even though neither group was statistically "wrong" in a way that suggested a total failure. In fact, for most of the writing categories, the differences in scores were small enough to be considered practically equivalent, meaning the machines were not wildly off-target. But for the content of the essays, the AI and the teachers diverged more noticeably, with the machines often rating the substance of the writing differently than the humans did.

The true story, however, emerged not from the numbers but from the words the teachers and the machines wrote in their feedback. The human teachers focused intensely on specific errors. Their comments were direct and corrective, pointing out exactly where a sentence was grammatically incorrect, where a paragraph lacked a clear topic, or where a specific word was misspelled. They acted like a coach spotting a flaw in a player's form, offering a precise fix. The AI systems, by contrast, took a much broader approach. They tended to cover all four areas of writing in a single response, often balancing praise with criticism. An AI might tell a student that their ideas were strong and well-organized before gently noting that the grammar needed work. They provided a wide net of observations rather than a targeted strike on specific mistakes.

This difference in approach led the researchers to propose a new way of thinking about how these tools should be used. The study suggests that artificial intelligence should not be viewed as a replacement for the teacher, nor as a machine that needs to be trained to think exactly like a human. Instead, the two offer different strengths that work best when combined. The AI can act as a tireless first reader, providing immediate, comprehensive feedback that covers every aspect of the writing and ensuring no student is left waiting for a response. The teacher then acts as the expert editor, reviewing that broad feedback, prioritizing the most important issues for the specific student, and offering the deep, contextual guidance that only a human can provide.

The researchers call this an evidence-informed framework for complementarity. It acknowledges that while the machines are excellent at scanning for patterns and generating consistent, multi-dimensional feedback, they lack the pedagogical judgment to know which feedback will actually help a specific learner at a specific moment. The human teacher brings the experience to decide what matters most. The study concludes that the future of writing assessment lies not in choosing between the human and the machine, but in coordinating them. By letting the AI handle the breadth of the initial review and the teacher handle the depth of the instruction, schools can create a system where students receive both the speed of technology and the wisdom of human expertise.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →