← Latest papers
💻 computer science

Simple use of small and medium sized local LLMs for Q-matrix generation

This paper demonstrates that open-weight small and medium-sized local LLMs can effectively and efficiently automate Q-matrix generation for Cognitive Diagnosis Models, offering a privacy-preserving and resource-accessible alternative to costly expert curation and closed-source frontier models.

Original authors: Simas Stočkus

Published 2026-09-22
📖 5 min read🧠 Deep dive

Original authors: Simas Stočkus

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of education, understanding exactly what a student knows is far more complex than simply assigning a grade. To truly diagnose a learner's strengths and weaknesses, educators rely on a specialized map called a Q-matrix. This map acts as a bridge, connecting specific test questions to the precise mental skills required to solve them. For instance, a question about adding fractions is not just a math problem; it is a test of a specific cognitive ability, such as finding a common denominator. Traditionally, creating these maps has been a slow, expensive process reserved for human experts who must painstakingly analyze every single question and skill. This manual work is so time-consuming that it often limits how quickly schools can adapt their teaching to individual needs. Furthermore, the modern push to use artificial intelligence to speed up this process has introduced new concerns. Many of the most powerful AI tools available today are massive, closed systems that require sending sensitive student data over the internet to distant servers, raising serious questions about privacy and security.

A new study by Simas Stočkus at Vilnius University explores a different path, asking whether smaller, more modest artificial intelligence models running directly on a standard laptop can perform this mapping task just as well. The researcher tested a variety of these local models, which are designed to be compact enough to run on consumer hardware without ever sending data outside the device. The goal was to see if these smaller models could accurately link test questions to the skills they measure, all while keeping student information private and avoiding the high costs of cloud computing. The study focused on a specific method of asking the AI to work: instead of asking the model to decide on one skill at a time for each question, the researchers asked the model to look at a question and list all the necessary skills in a single, complete answer. This approach mimics how a human teacher might glance at a problem and instantly recognize the full set of concepts involved, rather than checking them off one by one.

The results of this experiment were striking. When the researchers tested these local models across four different sets of educational data, ranging from fraction subtraction to probability theory, they found that the smaller models performed with remarkable accuracy. One particularly efficient model, known as Gemma 4 E4B, which contains only four billion parameters, managed to generate these skill maps with a high degree of correctness in just over five minutes for a full set of tests. This compact model outperformed several larger, more complex architectures that required significantly more computing power and time. The study demonstrated that by using a specific prompting strategy—where the AI is given clear examples of how to categorize skills before it begins—the smaller models could avoid common errors, such as guessing that a skill was needed when it was not. In fact, the most successful small model achieved a performance score of 66.02 percent, a result that rivals the best outcomes from much larger systems.

The research also highlighted a critical flaw in how some AI systems were previously being tested. When the models were asked to evaluate each skill in isolation, they often hallucinated, or invented, connections that did not exist, leading to inaccurate maps. However, when the models were allowed to see the entire list of skills at once and make a single decision for the whole question, their accuracy improved significantly. This shift in method proved that the models could understand the context of a question much better when they were not forced to work in a fragmented way. The study explicitly ruled out the idea that only massive, cloud-based supercomputers are capable of this task. Instead, it showed that open-source models, which can be downloaded and run on a personal computer, are a practical and privacy-safe alternative. These local models do not require the massive data centers or the expensive subscription fees associated with commercial AI services, making advanced educational analysis accessible to schools and researchers who cannot afford enterprise-level infrastructure.

Ultimately, this work suggests that the future of automated educational assessment does not necessarily depend on building ever-larger artificial intelligence systems. By refining how we ask these models to think and utilizing smaller, local versions, it is possible to create accurate diagnostic tools that respect student privacy and operate efficiently on everyday hardware. The study provides a clear roadmap for educators and developers who wish to implement cognitive diagnosis without compromising data security or breaking the bank. While the research focused on specific datasets and model types, the findings offer a strong foundation for future work, including the potential to fine-tune these models for specific subjects and to test them on real-world classroom assessments that have never been seen by an AI before. The evidence indicates that we can now build intelligent systems that are not only smart enough to understand student learning but also small enough to keep that learning safe.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →