On the retraining frequency of global models in retail demand forecasting
This paper demonstrates that less frequent retraining of global forecasting models across various machine learning and deep learning architectures maintains predictive accuracy while significantly reducing computational costs and energy consumption, offering a more sustainable and efficient approach to large-scale retail demand forecasting.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you run a massive grocery store with thousands of products. Every day, you need to guess how many apples, toothbrushes, or cereal boxes people will buy so you don't run out or waste money on unsold stock. To do this, you use a "Global Model"—a super-smart computer brain that learns from the sales history of all your products at once, rather than trying to learn a separate rule for every single item.
For a long time, the standard advice was: "Keep this computer brain constantly updated." Every time a new sale happens, you should retrain the model immediately to keep it sharp. The paper by Marco Zanotti challenges this habit. He asks: Do we really need to retrain this model every single day, or can we let it rest a bit?
Here is the breakdown of his findings using simple analogies:
1. The "Constant Training" vs. "Periodic Check-up"
Think of the forecasting model like a student studying for a big exam.
- The Old Way (Continuous Retraining): The student studies a little bit every single hour, every day. They are constantly re-reading the textbook the moment a new page is printed. This keeps them very up-to-date, but it's exhausting, expensive (in terms of time and energy), and they might get burned out.
- The New Way (Periodic Retraining): The student studies hard for a week, then takes a break. They only go back to the textbook once a month to update their notes with the latest chapters.
The Paper's Discovery: The study found that the student who studies once a month performs just as well (or sometimes even better) on the exam than the one who studies every hour. The "Global Model" is smart enough to learn the general patterns of the store (like "people buy more ice cream in summer") without needing to be re-taught every single day.
2. The Cost of "Green" Efficiency
Retraining these models isn't free. It requires massive computer power, which uses electricity and creates carbon emissions (like driving a car).
- The Analogy: Imagine the computer is a giant factory machine. Running it 24/7 to retrain the model every day costs a fortune in electricity.
- The Result: By switching from "daily updates" to "monthly updates," the study found that companies could cut their computing costs by 75%. If they stopped retraining entirely and just used the model once, they could save 90% of the cost.
- The Catch: For the most common type of prediction (guessing the exact number of items sold), the accuracy didn't drop at all. The model stayed sharp even while it "rested."
3. Machine Learning vs. Deep Learning
The paper tested two types of "brains":
- Machine Learning Models: These are like efficient, practical workers. They are great at following rules and patterns. The study found that these workers benefit the most from taking breaks. As the data gets bigger, these models save the most money and energy when retrained less often.
- Deep Learning Models: These are like genius artists who can see very complex, hidden patterns. They are powerful but require a lot of energy to run. While they also save money when retrained less often, they don't save quite as much as the practical workers, and they seem to hit a "ceiling" where taking longer breaks doesn't help them save much more time.
4. The "Safety Net" (Probabilistic Forecasting)
Sometimes, you don't just want to know how many apples will sell; you want to know the range of possibilities (e.g., "We will sell between 100 and 150 apples, but maybe up to 200"). This is called probabilistic forecasting.
- The Finding: If you need these "safety net" predictions, the model does need to be updated a little more often than for simple guesses. However, even here, updating it once every two weeks is usually enough. You don't need to do it daily.
The Bottom Line
The paper argues that the old rule of "always keep the model fresh" is outdated and wasteful.
- For Point Forecasts (Exact numbers): You can retrain your model once a month (or even less) and get the same accuracy as daily updates, while saving a massive amount of money and energy.
- For Safety Stock (Ranges): You might need to update it every two weeks, but still far less than daily.
In short: Your global forecasting model is like a well-trained athlete. It doesn't need to run a sprint every single hour to stay in shape. It can run a marathon, rest for a few weeks, and still win the race. By letting it rest, businesses can save money and help the environment without losing their edge in predicting demand.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.