Automated Scoring of Handwritten Mathematical Problem Posing Using Large Language Models
This study demonstrates that the large language model Gemini 3.5 Flash can achieve reliable automated scoring of handwritten middle school mathematical problem-posing tasks when using a zero-shot prompting approach with a structured rubric, particularly for whole number problems and when utilizing revised OCR text inputs.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a teacher with a mountain of homework to grade. You know that asking students to create their own math problems is a superpower for learning—it builds creativity and deep understanding. But grading these creative, handwritten answers is a nightmare. Every student writes differently, some make up impossible scenarios, and some just write nonsense. It takes forever, and even human teachers can disagree on what a "good" problem looks like. This is where Artificial Intelligence (AI) steps in, promising to be a tireless grading assistant. But here's the catch: AI is great at reading printed text, but can it actually understand messy, handwritten math scribbles and tell the difference between a brilliant student question and a confusing mess? This question sits at the intersection of education and computer science, exploring whether machines can learn to "think" like a math teacher when evaluating student creativity.
The researchers in this study decided to put a very smart AI, called Gemini 3.5 Flash, to the test. They wanted to see if this AI could read handwritten math problems created by 235 middle school students and grade them as accurately as human teachers do. To do this, they set up a "taste test" with two different scenarios: one involving whole numbers (like driving distances) and another involving fractions (like sharing a cake). They gave the AI two versions of the same handwritten text: one that was read directly by the computer's eyes (which sometimes makes mistakes with messy handwriting) and one that had been carefully corrected by human experts. They also tried giving the AI the grading rules in two different orders—starting from the lowest score and going up, or starting from the highest and going down—to see if the order of instructions changed the AI's mind.
The results were a mix of "great news" and "needs work." The study found that when the AI used the expert-corrected text, it agreed with human teachers much better than when it used the raw, messy computer reading. However, the type of math problem mattered a lot. The AI was much more accurate when grading problems about whole numbers (the driving task) compared to problems about fractions (the cake task). It seems the AI struggled more with the nuances of fractions, often giving high scores to answers that were mathematically correct but didn't actually fit the story the student was telling. Interestingly, the order in which the grading rules were presented to the AI didn't matter at all; whether the AI started from the bottom or the top of the rubric, its performance stayed the same.
In short, the study suggests that AI can be a reliable helper for grading handwritten math problems, but it isn't perfect yet. It works best when the handwriting is clear and the math involves whole numbers. The researchers found that simply feeding the AI better text (by having humans fix the computer's reading errors) made a noticeable difference, but the AI still needs help understanding the "story" behind the math, especially when fractions are involved. While the AI didn't get confused by the order of the instructions, it did get tripped up by the complexity of the math content. The authors conclude that while AI holds substantial potential to help teachers, we need to be careful about how we use it, especially for more complex topics like fractions, and we should always double-check that the computer is reading the student's handwriting correctly before it starts grading.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.