A Predictive Model to Detect Property Taxpayer Non-Compliance and Estimate Property Tax Liabilities in Rwanda Using Supervised Machine Learning
This study demonstrates that an XGBoost-based supervised machine learning framework, trained on over one million Rwanda Revenue Authority records, effectively detects property tax non-compliance and estimates liabilities, enabling a risk-based audit strategy that could recover 80% of unpaid taxes by targeting the top 20% of high-risk taxpayers.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In many parts of the world, collecting taxes on buildings and land is a difficult task. Governments need this money to build roads, schools, and hospitals, but the system often relies on people telling the truth about how much their property is worth. When owners declare a value that is too low, the government loses revenue, and honest taxpayers end up paying more than their fair share. To fix this, tax officials have traditionally relied on manual checks, where auditors physically visit properties or review paper files one by one. However, as cities grow and the number of buildings increases, this old way of working becomes too slow and expensive to keep up. In recent years, experts have begun to explore whether computers can learn from past records to spot patterns that humans might miss, helping officials decide who to check first without needing to look at every single file.
This is the challenge researchers from the University of Rwanda set out to solve. They wanted to know if a computer could look at the vast records held by the Rwanda Revenue Authority and figure out two things: first, which property owners are likely not paying the correct amount of tax, and second, what the correct amount should be. The team gathered a massive dataset containing information on more than one million property owners. This data included details about the buildings, how long the owners had been registered, how much they owed, what they had actually paid, and any penalties or interest they had accumulated. Instead of guessing, the researchers taught three different types of computer programs to learn from this history. They tested how well these programs could predict who was non-compliant and how accurately they could estimate the true tax bill for a property.
The researchers found that one specific type of computer program, known as XGBoost, was far better at the job than the others. This program is designed to learn by building many small decision steps and combining them to make a final prediction. When tested on data from 2024 that the computer had never seen before, this model correctly identified non-compliant taxpayers about 88 percent of the time. It was also very good at estimating the correct tax liability, with its predictions being much closer to the actual numbers than those made by the other models. The study revealed that the most important clues for spotting trouble were not always about the building itself, such as its size or location. Instead, the computer learned that a taxpayer's own behavior was the strongest signal. Factors like the ratio of what they paid versus what they owed, how much they owed in penalties, and how long they had been registered were the best indicators of whether someone was likely to be non-compliant.
The real power of this discovery lies in how it could change the way tax audits are conducted. The researchers ran a simulation to see what would happen if officials used these computer predictions to decide who to investigate. In a traditional approach, if officials picked taxpayers at random to audit, they would only recover about 20 percent of the unpaid money they were looking for. However, when they used the computer to rank taxpayers by risk and focused their efforts on the top 20 percent of the most likely offenders, the simulation showed they could recover roughly 80 percent of the unpaid liabilities. This suggests that by using data to guide their work, tax authorities could find the vast majority of missing revenue while spending far less time and money on audits.
The study concludes that this approach offers a practical path forward for Rwanda and other countries facing similar challenges. By integrating these predictive tools into their daily work, tax administrators could move away from random checks and toward a system where resources are directed exactly where they are needed most. The researchers emphasize that while the computer model is powerful, it should be used as a tool to support human decision-making, not to replace it entirely. They recommend that the Rwanda Revenue Authority adopt this risk-based system, while also continuing to improve the quality of their data and ensuring that the process remains fair. This work demonstrates that even in a resource-constrained environment, smart use of existing records can lead to a more efficient and equitable tax system.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.