← Latest papers
💻 computer science

An Optimized Stacking Ensemble Model for Forecasting Student Dropout

This study presents an optimized stacking ensemble model integrating K-Nearest Neighbors and Bayesian Network classifiers, enhanced by hyperparameter tuning and SMOTE-based balancing, which achieves 85.9% accuracy and outperforms existing methods to provide an effective early-warning tool for forecasting student dropout.

Original authors: Jackson Kisuule, Anthony Bua

Published 2026-09-23
📖 4 min read☕ Coffee break read

Original authors: Jackson Kisuule, Anthony Bua

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Every year, universities around the world face a quiet but persistent crisis: students who start their degrees but never finish them. This attrition is not just a personal tragedy for the individuals involved; it drains resources from institutions, lowers graduation rates, and leaves communities without the skilled professionals they need. For decades, schools have tried to solve this by manually reviewing grades and attendance, looking for students who seem to be slipping. However, these methods are often too slow to help a student before it is too late. In recent years, researchers have turned to machine learning, a branch of computer science where software learns to recognize patterns in data without being explicitly programmed with rules. By feeding historical records of thousands of students into these systems, computers can identify subtle warning signs that human observers might miss, offering a chance to intervene early and keep students on the path to graduation.

A team of researchers from Victoria University in Kampala, Uganda and The Unity University in Hargeisa, Somaliland, has taken this approach a step further by developing a new type of prediction system designed specifically for this challenge. Instead of relying on a single computer algorithm to make a judgment, they built a system that combines the strengths of two different methods. The first method, known as K-Nearest Neighbors, works by looking at a student's record and finding the most similar past students to see what happened to them. The second method, a Bayesian Network, calculates the likelihood of a student dropping out based on the complex web of relationships between their grades, financial status, and personal circumstances. While each method is useful on its own, the researchers found that they sometimes see different things in the same data. To solve this, they created a "stacking" model, a structure where a third, smarter program listens to the predictions of the first two and makes the final call. This final program acts like a seasoned editor, weighing the evidence from both sources to produce a more reliable verdict than either could achieve alone.

The researchers tested this system using a massive collection of real student records from The Unity University, covering the years 2021 through 2025. The dataset included over 4,400 students and tracked thirty-six different factors, ranging from how many courses a student passed to their financial standing and demographic background. A major hurdle in this type of research is that most students stay in school, while far fewer drop out, creating an uneven playing field for the computer. To fix this, the team used a technique to artificially balance the data, ensuring the system learned to recognize the warning signs of dropping out just as well as it learned to recognize students who were safe. They then split the data, using most of it to teach the system and holding back a portion to test how well it performed on students it had never seen before.

The results showed that the new combined system was indeed superior to the individual methods. When tested, the stacking model correctly identified whether a student would drop out or stay in school about 86 percent of the time. More importantly, it was very good at catching the students who were actually at risk. It correctly flagged 81 percent of the students who eventually left, a crucial metric for schools trying to intervene. In comparison, the single methods and a simpler way of combining them (called soft voting) were less accurate, missing more at-risk students or incorrectly flagging those who were safe. The analysis also revealed which factors mattered most: academic performance, specifically the number of approved courses and cumulative grades, was the strongest predictor, followed closely by financial indicators like tuition status and whether a student owed money to the institution.

This study demonstrates that by weaving together different types of machine learning, universities can build a more robust safety net for their students. The model does not replace human counselors but provides them with a powerful tool to prioritize their efforts. By identifying the students most likely to leave before they actually do, institutions can offer financial aid, academic tutoring, or counseling at the precise moment it is needed. While the researchers note that their model was trained on data from a single university and that future work should test it across different regions and with more advanced techniques, the findings offer a clear path forward. They prove that with the right combination of data and algorithms, it is possible to turn the tide on student attrition, turning a reactive process of managing dropouts into a proactive strategy for keeping students in school.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →