Cost-efficient generative AI summarization for scalable automated essay scoring in educational assessment
This study proposes a cost-efficient, hybrid generative AI framework that combines GPT-5-based essay summarization with handcrafted linguistic features to overcome transformer input limitations and enhance scalable automated essay scoring while balancing scoring reliability, summary fidelity, and computational cost.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a teacher with a mountain of essays to grade. You want to give every student helpful feedback, but reading thousands of long, rambling stories takes forever. This is the world of Automated Essay Scoring (AES), a branch of computer science where machines try to learn how to read and grade writing just like humans do. For a long time, these computers were like students who could only read short notes; if an essay was too long, the computer would just chop off the end, potentially missing the student's best arguments. But now, we have powerful "Generative AI" tools—super-smart computers that can write and summarize text. The big question researchers are asking is: Can we use these AI tools to shrink long essays into short summaries first, so the grading computer can read them easily, without losing the important parts? It's a bit like trying to fit a whole novel into a tweet; you need to keep the plot and the characters but cut out the fluff. If we get this right, schools could grade essays instantly, giving students feedback while they are still thinking about their ideas, rather than waiting weeks for a teacher to finish the pile.
This study dives into that exact problem, testing a new way to grade essays using a "summarize-then-grade" strategy. The researchers took a massive dataset of student essays and used three different versions of a powerful AI model to rewrite the long essays into shorter, 512-token summaries. They then fed these summaries, along with some original "handcrafted" clues about the writing (like sentence length and vocabulary), into a standard grading computer to see how well it matched human teachers' scores.
The results were a bit of a surprise. The most expensive, powerful model actually produced the best summaries—it kept the most details and meaning. However, when it came to the final goal of grading the essays accurately, the middle-sized model was the champion. It achieved the highest agreement with human raters, scoring a Quadratic Weighted Kappa (QWK) of 0.8435, which is a fancy way of saying it matched human judgment better than the others. The most expensive model scored 0.8350, and the smallest, cheapest model scored 0.8332. This suggests that for this specific job, you don't need the biggest, most expensive engine; the "mini" version offered the best balance of cost and performance.
However, the study also found a tricky pattern: the AI struggled more with the best essays. As the essay scores went up (from a 1 to a 6), the quality of the summaries dropped across all models. The researchers suggest this is because high-scoring essays are naturally longer and more complex, making them harder to shrink without losing important information. The "mini" model, while great overall, showed a bigger drop in performance on these complex essays compared to the full model.
The authors are careful to note that this isn't a final, perfect solution for every school in the world. They explicitly state that this is an "initial controlled evaluation," not a complete benchmark against every other possible method. They didn't test every other type of AI or every possible way to cut up the text. Instead, they showed that using a generative AI to summarize essays can work well and save money, but they also highlighted that we need more research to make sure the system doesn't accidentally penalize students who write long, complex, high-quality arguments. The study suggests that while we can use AI to make grading faster and cheaper, we have to be very careful about how we shrink the text so we don't lose the magic of a great essay in the process.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.