← Latest papers
⚡ electrical engineering

Evaluating Time Series Foundation Models for Electricity Price Forecasting: Contamination Risk, Distributional Shifts, and Covariate Dependence

This paper introduces a contamination-mitigated benchmarking framework to evaluate time series foundation models for electricity price forecasting, revealing that while these models are competitive and benefit from covariate support, they do not consistently outperform domain-specific methods, suggesting that ensembling both approaches yields the most promising results.

Original authors: Zhenghua Pan, Ahmed Aziz Ezzat

Published 2026-07-07
📖 4 min read☕ Coffee break read

Original authors: Zhenghua Pan, Ahmed Aziz Ezzat

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: The "Super-Genius" vs. The "Local Expert"

Imagine you are trying to predict the price of electricity for tomorrow. This is a tricky job because electricity prices are chaotic. They can be stable for weeks, then suddenly skyrocket or crash to zero due to weather, fuel shortages, or grid failures.

For a long time, people used Local Experts (specialized computer programs built just for electricity) to make these predictions. Recently, a new type of AI called a Time Series Foundation Model (TSFM) has arrived. Think of these as "Super-Geniuses." They have read almost every time-based dataset in existence (stock markets, weather, traffic, etc.) and can predict the future without ever being specifically taught about electricity.

This paper asks a simple question: Can the "Super-Genius" beat the "Local Expert" at predicting electricity prices, or do we still need the specialist?

The Problem: The "Cheating" Risk

There is a major catch with these Super-Geniuses. Because they are so smart, they might have already "read" the electricity data they are being tested on during their training. It's like giving a student a final exam that they already saw in their textbook. If they get a perfect score, is it because they are smart, or because they cheated?

The authors realized that many previous tests didn't check for this "cheating" (called contamination). To fix this, they created a Two-Test System:

  1. The Classic Test: A famous, old dataset everyone uses (GEFCom2014-P).
  2. The Fresh Test: A brand-new dataset (GridStatus2025) that was created after the Super-Geniuses finished their training. This ensures the AI hasn't seen these specific numbers before.

The Experiment: What They Tested

The researchers pitted the Super-Geniuses against three groups:

  1. The Old School: Simple math formulas (like guessing tomorrow's price is the same as today's).
  2. The General AI: Deep learning models trained on general data.
  3. The Local Experts: Models specifically designed by electricity engineers to understand how power grids work.

They also tested two versions of the Super-Geniuses:

  • The Blind Version: Only looks at past electricity prices.
  • The Informed Version: Can also look at "clues" like weather forecasts, fuel costs, and how much power people are using.

The Results: What Happened?

1. The Super-Geniuses are impressive, but not perfect.
On the "Fresh Test," the Super-Geniuses (especially the ones with access to clues) did very well. They crushed the simple math formulas and even beat many general AI models. However, they did not consistently beat the Local Experts. The Local Experts still hold the crown for the single best prediction.

2. The "Clues" are everything.
The paper found that the Super-Geniuses are like a detective who needs evidence. If you only let them look at past prices (the "Blind Version"), they struggle. But if you give them the "clues" (weather, fuel prices, load), they become much stronger. It turns out, knowing why the price might change is just as important as knowing what the price was yesterday.

3. The "Tail" Problem (The Extreme Events).
Electricity prices have "tails"—rare moments when prices go incredibly high or drop below zero.

  • The Local Experts are great at predicting these extreme swings.
  • The Super-Geniuses are okay, but sometimes they miss the mark on the very lowest or highest prices.
  • Analogy: If a storm is coming, the Local Expert knows exactly how bad the wind will be because they study storms. The Super-Genius knows storms happen, but might guess the wind speed is "average" when it's actually a hurricane.

4. The Magic Solution: The "Dream Team" Ensemble.
Here is the most exciting finding. When the researchers took the best Super-Genius and the best Local Expert and simply averaged their predictions together, the result was better than either one alone.

  • Analogy: Imagine a brilliant generalist doctor and a brilliant heart-specialist. If you ask them both for a diagnosis and take the average of their advice, you get a more accurate result than if you only listened to one. The paper suggests these two types of AI "speak different languages" and capture different pieces of the puzzle.

The Conclusion

The paper concludes that while the new "Super-Genius" AI models are powerful and competitive, they haven't completely replaced the need for specialized "Local Expert" models in the electricity world.

However, the best strategy isn't to choose one or the other. The future lies in combining them. By mixing the broad, general knowledge of the Foundation Models with the specific, deep knowledge of the electricity experts, we can create the most accurate and reliable forecasting system possible.

Key Takeaway: Don't fire the specialist just because you hired a genius. Instead, let them work together.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →