LMS-FAIR-India: A FAIR Compliant Synthetic LMS Interaction Dataset from a Private University for Privacy Preserving Learning Analytics
This paper introduces LMS-FAIR-India, a fully synthetic, FAIR-compliant dataset of 1.7 million LMS interaction events from a private Indian university that enables privacy-preserving learning analytics research without using real student data.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern classroom, learning rarely happens in isolation. It unfolds across digital platforms where students log in to watch lectures, download reading materials, submit assignments, and take quizzes. Every click, every pause, and every submission leaves a digital trace, creating a vast record of how people learn. Researchers call this field learning analytics, and they believe that by studying these patterns, they can understand what helps students succeed and what puts them at risk. However, a significant barrier stands in the way of this progress. The data that would reveal these patterns contains deeply personal information about real students. Laws designed to protect privacy, such as those in India and Europe, make it illegal to share this raw data openly. This creates a difficult situation: scientists want to study the data to improve education, but they cannot touch the data without violating the privacy of the individuals it describes.
To solve this dilemma, a researcher named Sanjay Agal from Parul University has created a new kind of resource. Instead of sharing real student records, which would be a breach of trust and law, the team generated a completely artificial dataset that mimics the behavior of real students. This collection, called LMS-FAIR-India, is a massive library of made-up interactions that look and feel exactly like the real thing but contain no actual people. It allows researchers to test their ideas and build better educational tools without ever needing to see a single real student's name or grade. The work demonstrates that it is possible to advance the science of learning while strictly honoring the right to privacy.
The core of this project is a synthetic dataset, which is essentially a computer-generated simulation of a university semester. The researchers built a model of a typical private university in India, complete with ten thousand virtual students, one hundred and fifty courses, and a six-month academic term. They did not use any real student records to create this model. Instead, they relied on public rules about how Indian universities operate, general statistics about how students typically behave online, and advice from faculty members to ensure the simulation felt authentic. The result is a collection of data that includes demographic details, course information, and over one point seven million recorded events, such as logins, video views, and quiz attempts. Every single piece of information in this dataset was created by a computer program following statistical rules, meaning there is zero risk of identifying a real person because no real person was ever involved in the process.
The dataset is structured to reflect the complexity of a real learning environment. It does not just list random events; it connects them logically. For instance, the data shows which students are enrolled in which courses, how long they spent on the platform, and what scores they received on their assignments. The researchers ensured that the patterns in this artificial data matched what is seen in real life. In the simulation, students who had stronger academic backgrounds before starting the semester tended to be more active on the platform, a pattern that mirrors reality. The data also captures the rhythm of student life, showing that activity peaks during the middle of the day on weekdays and drops significantly on weekends. By reproducing these natural behaviors, the dataset provides a realistic testing ground for researchers who want to develop tools to predict which students might struggle or to design better ways to engage learners.
One of the most important aspects of this work is how it handles the issue of privacy. In the past, researchers sometimes tried to protect real data by removing names or hiding specific details, a process known as anonymization. However, experts have shown that even anonymized data can sometimes be traced back to individuals if combined with other information. This new approach avoids that risk entirely. Because the data is generated from scratch using random distributions and statistical models, it is impossible to link any record back to a real student. The researchers tested this by trying to trick a computer into thinking a made-up student was real, and the computer failed to find any connection. This "privacy by design" method means the data can be shared freely with anyone in the world, allowing for open collaboration without the fear of legal or ethical violations.
To make the data useful for the global scientific community, the researchers followed a set of principles known as FAIR, which stands for Findable, Accessible, Interoperable, and Reusable. The dataset is hosted in a public online archive with a permanent digital address, so it can be easily found and cited. It is available for free under a license that allows anyone to use, modify, and share it, provided they give credit to the creators. The data is organized into clear, standard files that can be opened by almost any computer program, and it comes with a detailed guide that explains exactly what every column and number means. Furthermore, the researchers made the computer code used to create the data available to the public. This transparency allows other scientists to verify the results, adapt the simulation for their own universities, or generate new datasets tailored to different educational contexts.
The value of this dataset was proven through a series of tests designed to see if it could support real research. The researchers used the artificial data to train computer models to predict student engagement and to identify students who might drop out of their courses. The models performed just as well on this fake data as similar models have performed on real data in previous studies. For example, when the researchers tried to predict which students would finish their courses, the system correctly identified at-risk students with a high degree of accuracy. They also confirmed that the relationship between how often a student logs in and their final grades was preserved in the simulation, matching the findings from real-world studies. These results suggest that the synthetic data is a reliable substitute for real data when it comes to developing and testing new educational technologies.
This work addresses a critical gap in the field of learning analytics, particularly for the context of higher education in India. While there are some public datasets from Western countries, they often reflect different educational systems and cultural contexts. The LMS-FAIR-India dataset fills this void by providing a large-scale, realistic simulation of an Indian private university. It offers a unique opportunity for researchers to study how students in this specific environment learn and interact with technology. By providing a tool that is both privacy-safe and scientifically robust, the project enables a new wave of innovation in education. It allows scientists to ask difficult questions about student success and to find answers without compromising the dignity or safety of the individuals they are studying.
The implications of this work extend beyond just one dataset. It demonstrates a practical path forward for the entire field of educational research. By showing that high-quality, privacy-preserving data can be generated and shared openly, the researchers have removed a major obstacle that has slowed down progress for years. Other institutions can now use the same methods to create their own synthetic datasets, tailored to their specific needs, without needing to navigate complex legal hurdles. This approach fosters a culture of openness and collaboration, where the focus remains on improving learning outcomes for all students. The dataset stands as a testament to the idea that protecting privacy and advancing science are not opposing goals, but rather complementary aims that can be achieved together through careful and creative design.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.