Explainable Machine Learning for Early Detection of Mathematics Learning Difficulties: A Systematic Review of Multimodal Educational Data and Predictive Modeling
This systematic review synthesizes research on explainable machine learning methods utilizing multimodal educational data to enable the early detection and transparent explanation of mathematics learning difficulties, while highlighting current challenges and proposing a future research agenda focused on privacy, standardization, and ethical deployment.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Mathematics is a language of logic that most children learn to speak fluently, but for a small number of students, the numbers simply do not add up. This struggle, known as mathematics learning difficulty, affects roughly three to six percent of school-aged children. It is not merely a case of being slow at arithmetic; it often involves a fundamental disconnect with how numbers relate to one another, leading to a reliance on immature counting strategies long after peers have moved on to automatic recall. When left unaddressed, this difficulty can spiral into a lifetime of anxiety, reduced confidence, and missed opportunities in science and technology fields. For decades, schools have relied on standardized tests to identify these students, but these exams are like a rearview mirror: they only show what has already happened. By the time a test score confirms a student is falling behind, the student has often already experienced repeated failures, and the specific reasons for their struggle—whether it is a hidden misconception or a bad habit of guessing—remain invisible to the teacher.
A team of researchers has recently turned its attention to a different way of looking at the problem, one that happens while the student is still working. They conducted a systematic review of the latest scientific studies to see if modern technology could spot these difficulties much earlier, before the damage is done. Instead of waiting for a final grade, they examined how computers can analyze the tiny, moment-to-moment actions of a student solving a math problem. This approach gathers a wide variety of signals, from the clicks and time spent on a digital screen to the movement of a student's eyes as they scan a page. The researchers wanted to know if combining these different streams of information, and then using computer programs that can explain their own reasoning, could create a system that is both accurate and trustworthy enough for teachers to use in a real classroom.
The researchers sifted through thousands of scientific papers, narrowing their focus to eighteen studies that met strict criteria for quality and relevance. What they found was a landscape dominated by two main types of data. The most common source was the digital trail left behind by students using online learning tools. These logs record every click, every hint requested, and every second spent on a problem. By analyzing these patterns, computers can identify a specific behavior called "wheel-spinning," where a student keeps trying the same problem over and over without ever making progress, or conversely, giving up too quickly. The second major source of data was eye-tracking technology, which watches where a student looks while they solve a problem. This method revealed something the logs could not: the specific way a student's eyes move can expose a deep misunderstanding of a concept, such as how they estimate the size of a number, often within the first fifteen to forty-five seconds of starting a task.
While these technologies showed promise, the researchers discovered that the real breakthrough lies in how the data is combined and explained. In the past, computer models that predicted student struggles were often "black boxes," meaning they could give a warning but could not say why. The studies reviewed in this paper increasingly use "explainable" methods, which act like a transparent window into the computer's thinking. These systems can point to specific behaviors, such as "this student is guessing because they are answering too fast," or "this student is stuck because they are looking at the wrong part of the diagram." The researchers found that when these explainable models are fed with multiple types of data at once—like combining the timing of clicks with the path of the eyes—they become much better at spotting trouble early. One study noted that these combined approaches could distinguish between different types of learning difficulties with high accuracy, sometimes within the first few attempts at a problem.
However, the path from a computer screen to a helpful classroom intervention is not without its hurdles. The review highlighted a significant gap in the current research: most of these systems are trained on "proxy" labels rather than official medical diagnoses. In simpler terms, the computers are often taught to recognize "struggle" based on low test scores or high error rates, rather than being taught to recognize the specific clinical condition of developmental dyscalculia. This means the models might be identifying students who are having a hard time for many different reasons, not just the specific learning difficulty the researchers hope to catch. Furthermore, the researchers warned that these tools are not yet ready to be used universally. A system that works perfectly on one online math platform might fail completely on another, and there are serious concerns about whether the data might unfairly penalize students from certain backgrounds or with different learning styles.
The ultimate goal of this research is not to replace teachers with machines, but to give educators a new kind of insight. The authors propose a four-step process where a system first detects a risk, then diagnoses the specific nature of the problem, suggests an intervention, and finally monitors whether that help is working. For this to succeed, the technology must be transparent enough that a teacher can understand the warning and trust it. The review concludes that while the potential for early detection is real and the technology is advancing rapidly, the field needs to move beyond just building accurate models. Future work must focus on creating shared, privacy-safe datasets that allow these systems to be tested across different schools, and it must involve teachers directly in the design process to ensure the warnings they receive are practical, fair, and truly helpful for the students sitting in front of them.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.