Cost-Aware Evaluation of Time-Series Foundation Models for Urban Air-Quality and Temperature Forecasting
This study demonstrates that for urban air-quality and temperature forecasting in resource-constrained cities, zero-shot foundation models often match or outperform specialized and transfer-learning approaches when evaluated under realistic data and compute constraints, challenging the assumption that heavy foundation models are unsuitable for such deployments.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Cities are breathing in a new kind of danger, one that requires constant, hyper-local monitoring to keep people safe. Fine particles in the air, invisible to the eye but deadly to the lungs, are a leading cause of premature death worldwide, hitting the hardest in places where the air is most polluted and the data is most scarce. To fight this, city planners and health officials need to know what the air will do in the next few hours, right at the level of a single street corner. They also need to predict local temperatures to manage heatwaves and energy use. The problem is that the cities needing these forecasts the most often lack the long, clean records of data and the expensive computer power required to build custom prediction tools. For years, the standard advice for these resource-strapped operators has been to build tiny, efficient models from scratch or to borrow knowledge from data-rich cities, operating under the assumption that the newest, most powerful prediction engines are simply too heavy and complex to run on the edge of the network.
A team of researchers from the University of Rajshahi in Bangladesh decided to test this long-held assumption. They set out to see if the latest generation of "foundation models"—massive artificial intelligence systems pre-trained on vast amounts of data from around the world—could actually forecast air quality and temperature for a specific city without ever being trained on that city's data. They compared three different approaches: training a small, custom model locally; borrowing a model from a data-rich city and tweaking it for the local conditions; and simply downloading a pre-trained foundation model and letting it work immediately, a method known as zero-shot learning. They tested these strategies across twenty-nine cities, ranging from data-rich hubs like Seoul and New York to data-scarce locations like Nairobi and Kampala, looking at both hourly air pollution levels and local temperatures.
The results overturned the conventional wisdom that foundation models are too heavy for local use. When forecasting air pollution, the pre-trained foundation model performed just as well as the carefully tuned local specialist, even though the foundation model had never seen a single data point from the local city. In fact, when the researchers accounted for the cost of energy and the time required to train models, the pre-trained option often came out ahead. The foundation model used roughly ten times less energy than the local models when the system needed to be retrained frequently, a common requirement for keeping forecasts accurate. The only time the local, custom-trained model became more energy-efficient was if it was trained once and then left alone for nearly a year, a scenario that rarely happens in the fast-changing world of air quality monitoring.
The study also uncovered a hidden flaw in how many scientific comparisons are made. When predicting temperature, the local specialist model appeared to be vastly superior, but this advantage vanished once the researchers corrected for a subtle timing error. The original tests had given the local model a "perfect weather forecast" of the future conditions it needed to make its prediction, an unfair advantage that does not exist in the real world. Once the researchers restricted the model to using only the weather data actually available at the moment the forecast was made, the local model's lead shrank to a negligible difference, and the pre-trained foundation model remained a strong, reliable contender. This finding suggests that many previous studies claiming one method is better than another may have been misled by the same timing error.
Ultimately, the researchers found that for cities with limited data and tight budgets, the best strategy is often to stop trying to build a custom engine and start downloading a pre-made one. The pre-trained foundation model proved to be statistically indistinguishable from the best local models in accuracy, while offering a massive advantage in speed and energy efficiency. It works without needing a local team of engineers or a long history of local data. The study concludes that the practical default for forecasting in data-scarce cities has shifted: instead of training a small model, the most effective and efficient path is to download a large, pre-trained one and let it do the work. This change could make high-quality air quality and temperature forecasting accessible to cities that have long been left behind, turning a complex engineering challenge into a simple matter of downloading a file.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.