Atom Learning Model (ALM): how a real classroom got tokenised
This paper describes the Atom Learning Model (ALM), a system that tokenized two mathematics textbooks into a prerequisite graph of 1,934 learning steps to generate personalized questions for 373 students at a low cost, while revealing unexpected findings such as the high cost of establishing links, the inability of language models to predict question difficulty, and the fact that the system's deployment stopped just short of testing its core premise.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In a typical classroom, a teacher stands before thirty students, all roughly the same age, and tries to teach a single lesson. The challenge is that no two children in that room are at exactly the same place in their learning. One might be ready to solve a complex puzzle, while another is still struggling with the basic pieces. Because the teacher cannot instantly know the precise skill level of every single student, the lesson is pitched at a middle ground, leaving some bored and others lost. For decades, educators have tried to solve this by sorting students into groups based on their general ability, but this only narrows the gap slightly; it still treats thirty children as one unit rather than thirty individuals. The core problem is a lack of resolution: schools have a map of what needs to be learned, but they lack a way to pinpoint exactly where each child stands on that map at any given moment.
This is the gap that a new system called the Atom Learning Model attempts to fill. The idea is to break a school subject, like mathematics, down into its smallest possible steps, called "atoms." Imagine a skill like solving an equation not as a single, monolithic task, but as a chain of tiny, distinct actions: knowing how to multiply, knowing how to subtract, knowing how to rearrange a line. If a student can do the first three but fails the fourth, the system knows exactly where the break in the chain is. By mapping these tiny steps and the order in which they must be learned, the system can theoretically hand every child a worksheet that is perfectly pitched just beyond what they already know. This approach promises to move education away from rigid year groups and toward a model where students progress at their own pace, mastering one small step before moving to the next.
A researcher at Imperial College London, led by Philipp Bogdan, tested this idea in a real-world setting over seven weeks in two secondary schools in England. They did not write the curriculum or the questions by hand. Instead, they fed two standard mathematics textbooks into a computer system and asked it to read the pages, extract the skills, and build a massive, interconnected map of knowledge. The system processed 757 pages of text and turned them into 1,934 distinct "atoms," each representing a single thing a learner can do. It then spent the majority of its effort figuring out how these atoms connect, creating over 4,600 links that show which skills must be learned before others. This structure allowed the computer to generate thousands of unique questions, tailored to specific combinations of these tiny skills, and serve them to 373 students.
The results of this experiment were surprising and challenged several common assumptions about how learning works. First, the researchers found that a computer model looking at a question could not predict how hard it would be for a student. The system had assigned difficulty labels to the questions, but when they compared these labels to how students actually performed, the labels were completely useless. The difficulty of a question, they discovered, is not a fixed property of the question itself. Instead, difficulty is a relationship between the question and the specific student answering it. A question is only hard if the student is missing the specific tiny steps required to solve it. This means that a teacher or a computer cannot simply look at a problem and say it is "hard"; they can only know it is hard for a particular child who lacks the necessary background steps.
Another unexpected finding concerned the speed of feedback. The researchers measured how long students waited for their answers to be graded and how that wait time affected their willingness to keep working. They found that if a student had to wait just seven seconds for a mark, they were significantly less likely to start the next question compared to if they waited only three seconds. This tiny delay, which most educators would consider negligible, was enough to break the student's momentum. It suggests that in a digital learning environment, the speed of the system is just as critical to engagement as the quality of the questions.
Perhaps the most significant limitation of the study was that the system never actually tested its own deepest theory. The central premise of the Atom Learning Model is that difficulty comes from the depth of the chain of skills a student is missing. To prove this, the system would need to serve questions that require a student to bridge a gap of three or four missing steps. However, in the seven-week trial, the system almost never served a question that went deeper than two steps. The researcher admitted that the most important part of their theory—the idea that difficulty is purely about what is missing—remained untested because the system never asked the hard questions that would have proven it. The system worked well for shallow questions, but the ultimate test of whether it can guide a student through a deep gap in knowledge was left for the future.
Despite these limitations, the experiment provided a concrete look at what a fully tokenized curriculum could cost and how it might function. The entire process of reading the textbooks and building the map of skills cost between £615 and £1,230, a one-time investment. Once built, the system generated over 6,000 questions for the students at a cost of about 26 pence per question. The study showed that it is possible to automate the creation of a personalized learning path without human teachers writing every single problem. However, the researcher was careful to note that the system did not yet replace the teacher. The map of skills existed, and the questions were generated, but the system did not yet use the student's live performance data to decide what to show them next in real-time. The scores were calculated after the fact, meaning the students were not actually being guided by the system's full potential during the trial.
The vision that emerges from this work is one where the concept of a "year group" becomes obsolete. If a system knows exactly which tiny step a child has mastered and which one they are missing, there is no need to keep thirty children together in a single room waiting for the rest of the class to catch up. A child could move through the material at their own speed, spending nine months on a level if they need to, or moving through it in six months if they are ready. The exam, currently a logistical event where everyone sits down at the same time, could become a personal milestone that a student takes whenever they are ready. The Atom Learning Model did not achieve this future in the seven-week trial, but it built the first real bridge toward it, showing that the plumbing for such a system can be constructed, even if the water has not yet flowed through it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.