Demand Forecasting for Internet Bandwidth using Machine Learning to Enhance Resource Management in Nepal
This research proposes the development and deployment of advanced machine learning models to accurately forecast internet bandwidth demand in Nepal's Kathmandu Valley, enabling ISPs to optimize resource allocation, mitigate network congestion, and improve overall service quality through data-driven decision-making.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The internet is the invisible infrastructure that powers modern life, connecting people to information, business, and each other. Behind the scenes, Internet Service Providers act as the gatekeepers of this flow, managing the massive streams of data that travel through their networks. However, these networks have a physical limit to how much information they can carry at any single moment. When too many people try to use the internet at once, the pipes become clogged, leading to slow speeds, dropped connections, and a frustrating experience for everyone. The challenge for providers is to predict exactly when and how much data will be needed so they can prepare their networks in advance. If they guess too low, the network crashes; if they guess too high, they waste money on expensive equipment that sits idle. To solve this, researchers are turning to machine learning, a branch of computer science where software learns from past patterns to make better guesses about the future, rather than relying on simple averages or rigid rules.
In the Kathmandu Valley of Nepal, where internet demand is growing rapidly, a team of researchers set out to find the most reliable way to forecast this bandwidth usage. They focused on a specific question: which method can best predict how much data users will consume on any given day? To answer this, they worked with eighteen months of real-world data from Dish Media Network, a major provider in the region. This data was a detailed log of every connection, recording exactly when users logged on, how long they stayed connected, and how much information they downloaded. The researchers cleaned this massive dataset, filling in gaps where logs were missing and organizing the information by day, month, and hour to see the big picture of how the network was being used.
The team tested three different approaches to see which one could make the most accurate predictions. The first two methods were traditional statistical tools often used for forecasting. One was designed to spot simple trends, while the other added a layer to account for repeating seasonal patterns, such as higher usage during certain times of the year. The third method was a more advanced type of artificial intelligence known as a Long Short-Term Memory model. Unlike the traditional tools, which look for straight lines and regular cycles, this model is built to understand complex, shifting behaviors in data, much like how a human might learn to anticipate a friend's habits by watching them over a long period.
When the researchers compared the results, the differences were clear. The traditional methods provided a baseline, but they struggled to capture the full complexity of how people actually used the internet. The model that accounted for seasons performed slightly worse than the simpler trend model, suggesting that the seasonal patterns in this specific network were not strong enough to improve the forecast. The advanced artificial intelligence model, however, outperformed both. It produced the smallest gap between what it predicted and what actually happened, with an error rate significantly lower than the other two. This indicated that the neural network was better at recognizing the subtle, non-linear ways that user behavior changes over time.
The study also uncovered what actually drives these changes in usage. The number of active users was the single most important factor, showing a strong link to how much data was consumed. The time of day also played a crucial role, with usage spiking between late morning and early afternoon, while the specific day of the week mattered far less. These findings suggest that for providers in this region, focusing on the total number of users and the time of day is more valuable than trying to predict based on the calendar date alone.
Beyond the technical results, the researchers emphasized the importance of using these tools responsibly. They noted that for such systems to be fair and trustworthy, the data used to train them must represent a wide variety of people and regions. If the data is biased, the predictions could lead to unfair service decisions. The authors argue that with careful oversight, clear documentation, and regular checks for bias, these powerful tools can help providers manage their networks efficiently without compromising the quality of service for their customers. The study concludes that while older methods have their place, the more sophisticated machine learning approach offers a significant step forward in keeping the internet running smoothly for everyone.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.