Attention Without Estimation: Zero-Shot Time Series Foundation Models versus Long-Memory Econometrics in Realized Volatility Forecasting
This paper establishes a rigorous theoretical and empirical framework comparing zero-shot time series foundation models against classical long-memory econometric benchmarks for realized volatility forecasting, finding that while foundation models outperform GARCH and EWMA, they ultimately compete closely with HAR-family models that capture most of the exploitable signal in volatility data.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to predict how bumpy a car ride will be tomorrow. In the world of finance, this "bumpiness" is called volatility, and getting it right is crucial for pricing options, managing risk, and keeping banks safe. For decades, the experts have used a specific set of math tools called econometrics (think of them as seasoned, old-school mechanics) to forecast this. Recently, a new challenger has entered the garage: Time Series Foundation Models (TSFMs). These are giant, super-smart AI brains trained on millions of different time-series patterns, promising to predict volatility for any asset without needing to be re-tuned for each specific car.
This paper is a rigorous, no-nonsense race between these two approaches. Here is what happened when they hit the track.
The Race Track: A Synthetic World
First, a crucial detail: the authors didn't test this on real-world stock markets (which are messy and full of secrets). Instead, they built a perfectly controlled, synthetic race track using a computer simulation. They created 20 fictional assets (like 20 different cars) that moved according to a strict set of rules involving random jumps and smooth curves.
- 15 of these cars were used to "train" the AI brain (the TSFM).
- The other 5 cars were kept completely secret. The AI had never seen them before. This is called zero-shot testing—like asking a chef who has only cooked Italian food to make a perfect Thai dish without a recipe.
The Contenders
The Old Guard (Econometrics):
- GARCH & EWMA: These are the basic, reliable mechanics. They look at recent bumps and guess the next one. They are simple but sometimes miss the big picture.
- HAR-RV: This is the "star mechanic." It doesn't just look at yesterday; it looks at the last day, the last week, and the last month all at once. It's famous for being incredibly hard to beat.
- HARQ & HAR-CJ: These are the "super-mechanics." They are special versions of the star mechanic that can specifically handle measurement noise (when the speedometer is glitchy) and sudden jumps (when a car hits a pothole).
The New Challenger (TSFM-Surrogate):
- This is the AI brain. It uses a "patch-attention" mechanism (imagine it glancing at chunks of the road at once rather than every single inch). It was trained on the 15 known cars and then asked to predict the bumps for the 5 secret cars without any extra tuning.
The Results: Who Won?
1. The AI vs. The Basic Mechanics
The AI brain (TSFM-surrogate) crushed the basic mechanics (GARCH and EWMA).
- In the tests, the AI made significantly fewer errors. It was statistically proven to be better at predicting the bumps than the old-school models.
- The Verdict: The AI is definitely better than the basic tools.
2. The AI vs. The Super-Mechanics (The Twist)
Here is where it gets interesting. When the AI raced against the HAR family (HAR-RV, HARQ, and HAR-CJ), the results were a tie in the big picture, but the super-mechanics won the head-to-head.
- Head-to-Head: The HAR models (especially the ones that handled jumps and noise) were statistically better than the AI. The AI made more mistakes than the HAR models when predicting the exact size of the bumps.
- The Big Picture (The Model Confidence Set): However, when the judges looked at the whole group together, they couldn't say for sure that the HAR models were definitely better than the AI. The AI was statistically indistinguishable from the best HAR models.
- The Takeaway: The AI is good enough to be in the "Hall of Fame" with the best econometric models, but it hasn't quite beaten them yet.
3. The "Jump" Factor
The simulation included sudden "jumps" (potholes). The HAR models that were specifically designed to spot these jumps (HAR-CJ) and handle noisy data (HARQ) performed the best. The AI, which wasn't explicitly taught to look for jumps, didn't do as well as these specialized models. This suggests the AI needs more data or better training to learn these specific patterns on its own.
Why Didn't the AI Win?
The authors suggest a few reasons, based on their math and simulations:
- Data Hunger: The AI was only trained on 15 assets. The paper argues that foundation models usually need thousands or millions of assets to truly shine. With only 15, the AI is still learning the ropes, while the HAR models (which are simpler) can learn perfectly from just one asset's history.
- Context Window: The AI looked at the last 32 days of data. The authors found that when they let the AI look at 64 days, it got better. This suggests the AI just needs a longer "lookback" to see the full pattern, similar to how the HAR models look back a full month.
- Specialization: The HAR models are like specialized tools built specifically for volatility. The AI is a "generalist" trying to learn everything at once. In a controlled, single-regime simulation, the specialist wins.
What the Paper Explicitly Rules Out
- It is NOT a win for AI over all econometrics. The paper explicitly states that the HAR family remains the hardest benchmark to beat.
- It is NOT a test of real-world models. The AI used here is a "surrogate" (a stand-in) built from scratch. The authors did not test famous commercial models like Chronos or TimesFM because they couldn't download them in their offline environment. They are careful to say their results apply to the paradigm of foundation models, not necessarily those specific brands.
- It is NOT a solved problem. The paper does not claim AI has replaced the need for econometrics. In fact, it suggests that for a single asset in a stable market, the old-school math is still very hard to beat.
The Bottom Line
In this simulated race, the Time Series Foundation Model proved it is a serious contender, beating the basic tools and keeping pace with the best experts. However, it hasn't dethroned the HAR family yet.
The paper suggests that the AI's true superpower isn't beating the experts in a small, controlled room, but rather its ability to work zero-shot (without retraining) across many different, messy, real-world markets at once. If you have thousands of assets and sudden market crashes (regime shifts), the AI might eventually pull ahead. But for now, in this specific simulation, the long-memory, jump-aware econometric models are still the kings of the hill.
The authors conclude that we need more data (more assets) and longer lookback windows to see if the AI can truly surpass the specialists. Until then, the best strategy might be to keep the old-school math as a safety net while the AI learns the ropes.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.