Intrinsic predictability rather than urban structure governs machine learning forecast skill for electricity demand
This paper argues that the forecast skill of machine learning models for urban electricity demand is primarily determined by the intrinsic predictability of a city's historical load patterns (specifically seasonal variance and temperature sensitivity) rather than by the city's structural characteristics or geographic location.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Cities are increasingly turning to artificial intelligence to manage their electricity needs, hoping that smart algorithms can predict exactly how much power people will use next month. This promise of precision is central to modern urban planning, where officials must balance the grid, cut costs, and reduce emissions. The logic seems straightforward: if a computer can forecast demand accurately, it proves the city's energy system is well understood; if the forecast fails, it suggests something hidden or broken is happening within that specific city's infrastructure or behavior. This assumption has driven a wave of investment in machine learning tools for public administration, treating prediction errors as direct clues to local problems that need fixing. However, this approach rests on a critical, often unexamined belief: that the ability of a model to predict the future depends on the unique structure of the city itself.
A new study challenges this fundamental belief by looking at electricity demand across forty-six large cities in the United States. The researchers, led by Mahyar Hassani and colleagues, did not simply build a better forecasting engine; they built a diagnostic tool to test whether the engine's success or failure actually tells us anything about the city. They analyzed twelve years of monthly electricity data, from 2007 to 2018, applying a rigorous testing method that prevents the computer from using future data. Their goal was to see if the accuracy of the predictions varied because of how big a city was, how wealthy it was, or how its buildings were arranged, or if the variation came from something else entirely.
The results were surprising. The study found that the skill of the machine learning model had almost nothing to do with the city's size, its population, or its specific urban layout. A city with a million people was not inherently easier or harder to predict than a city with a quarter of a million. Instead, the ability to forecast demand was governed almost entirely by the city's own history of electricity use, specifically two simple patterns: how strongly the demand followed the seasons and how tightly it tracked the temperature. Cities with a strong, predictable rhythm of heating in winter and cooling in summer were easy to forecast, regardless of whether they were in the desert or the Midwest. Cities with milder climates, where the weather and electricity use did not swing dramatically from month to month, were much harder to predict. The researchers found that these two factors—the seasonal rhythm and the temperature sensitivity—explained nearly all the differences in forecast accuracy between the cities.
To reach this conclusion, the team had to strip away several layers of potential confusion. They compared their advanced machine learning models against a very simple rule: guessing that next month's electricity use would be the same as the same month last year. In nearly half of the cities, this simple "seasonal-naïve" rule performed just as well as the complex artificial intelligence. In some cases, the simple rule was actually better. This showed that for many cities, the "signal" of the seasons was so strong that no amount of computer power could add much value. For the cities where the complex models did better, the improvement was often marginal. The study demonstrated that when a model fails to predict a city's demand, it is rarely because the city is chaotic or poorly governed. Instead, the failure is usually because the city's electricity use simply lacks a strong, predictable pattern to begin with.
The researchers also uncovered a hidden danger in how data is handled. They showed that a common, routine practice of filling in missing data points with a repeated value could create a fake geographic pattern. If a few cities happened to have missing data that was filled in this way, the models would suddenly appear to fail in a specific region, creating the illusion of a regional problem where none existed. This finding serves as a stark warning: what looks like a deep insight into urban governance might just be an artifact of how the numbers were cleaned. By controlling for these data quirks and the intrinsic predictability of the demand, the study found no evidence that forecast skill clusters by region or that it reveals hidden local failures.
The practical takeaway for city officials is a shift in how they interpret these forecasts. A poor prediction is not, on its own, evidence of a broken system or a hidden local crisis worth investigating. It becomes evidence of a problem only after the city's inherent predictability has been accounted for. If a city has a flat, unseasonal demand pattern, a low forecast score is expected and tells the official nothing new. If a city has a strong seasonal pattern but the model still fails, that is when the alarm should ring. The study suggests that before spending money on audits or new programs to fix "bad" forecasts, administrators should first check if the city's demand history actually allows for a good forecast. By focusing on the intrinsic nature of the data rather than the complexity of the algorithm, public officials can avoid chasing ghosts and focus their resources on the cities and systems where real problems actually exist.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.