← Latest papers
📊 statistics

Fidel-TS: A High-Fidelity Multimodal Benchmark for Time Series Forecasting

To address the limitations of existing benchmarks caused by data contamination and leakage, the authors introduce Fidel-TS, a new high-fidelity, large-scale multimodal benchmark designed to provide more accurate and reliable evaluations of time series forecasting models.

Original authors: Zhijian Xu, Wanxu Cai, Xilin Dai, Zhaorong Deng, Qiang Xu

Published 2026-05-11
📖 5 min read🧠 Deep dive

Original authors: Zhijian Xu, Wanxu Cai, Xilin Dai, Zhaorong Deng, Qiang Xu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to judge who is the best chef in the world. To do this fairly, you need a kitchen with fresh ingredients, a clear recipe, and a way to test them without letting them peek at the answers beforehand.

For years, the field of Time Series Forecasting (predicting future trends based on past data, like stock prices or weather) has been using "old, stale recipes" to test its models. The paper Fidel-TS argues that these old tests are broken, leading us to believe some chefs are geniuses when they might just be cheating or lucky.

Here is the paper explained in simple terms, using analogies:

1. The Problem: The "Cheating" Kitchen

The authors say current benchmarks (the tests used to grade AI models) have three major flaws:

  • The "Stale Menu" (Data Contamination): Many datasets used for testing are years old. Because they are so old and famous, the AI models (especially the big "Large Language Models" or LLMs) have likely already eaten them during their training. It's like giving a student a math test they've already memorized the answers to. They get a perfect score, but it doesn't prove they actually know math.
  • The "Leaky Window" (Data Leakage): In previous tests, researchers would go back in time and grab text (like news articles) that happened after the event they were trying to predict. It's like asking a weather forecaster to predict tomorrow's rain, but secretly handing them a newspaper from tomorrow that says "It rained." The model isn't predicting; it's just reading the answer key.
  • The "Messy Counter" (Structural Confusion): Old tests mixed up different types of data. They didn't clearly distinguish between "different sensors in the same building" versus "the same sensor in different buildings." This made it hard to tell if a model was truly smart or just memorized a specific location.

2. The Solution: Fidel-TS (The "High-Fidelity" Kitchen)

To fix this, the authors built a new benchmark called Fidel-TS. Think of this as a brand-new, ultra-modern kitchen designed with strict rules to ensure fairness.

  • Fresh Ingredients (API Streams): Instead of using old, static files, they pull data directly from live, secure internet connections (APIs). This ensures the data is fresh, high-frequency (happening every few minutes), and, crucially, not in the training data of the AI models. No cheating allowed.
  • The "Scheduled Menu" (Leak-Free Text): When they add text (like weather reports) to help the models, they only use things that are scheduled in advance. For example, a weather forecast is released at a specific time before the weather happens. They avoid random news articles that might accidentally reveal the future. This is like giving the chef a menu of ingredients that will arrive tomorrow, rather than a list of what did arrive.
  • Clear Stations (Subjects vs. Channels): They organized the data clearly.
    • Subjects: The specific location (e.g., "Sensor A on 5th Street").
    • Channels: The type of data (e.g., "Traffic Speed").
      This allows them to test if a model can learn from one street and apply that knowledge to a new street it has never seen before (Generalization).

3. The Taste Test: What They Found

The authors cooked up a massive experiment, testing various AI chefs (models) in this new, fair kitchen. Here is what they discovered:

  • The "Specialist" Chefs are Still Best: In the old tests, big "Foundation Models" (general-purpose AIs) claimed they could predict anything without extra training. In this new, strict kitchen, they struggled. They performed well on short-term predictions but fell apart when asked to look further ahead. The specialized models (chefs who only cook time series) still won.
  • Reading the Menu Helps (Sometimes): When models were allowed to read the "scheduled" text (like weather forecasts), some got better. But it depended entirely on the model's architecture. Some models ignored the text; others used it to spot sudden changes (like a storm causing a traffic jam) that pure math couldn't see.
  • The "Generalist" Chefs (LLMs) Struggle: The big, general-purpose AI models (like the ones you chat with) were surprisingly bad at this specific task when tested fairly.
    • They made many mistakes in precision.
    • They often failed to follow instructions (like formatting their answer correctly).
    • When the task got harder (longer predictions or more variables), they crashed.
    • The Takeaway: Just because an AI is good at writing poems or coding doesn't mean it's good at predicting the future.

4. Why This Matters

The paper concludes that we have been overestimating the progress of AI in time series forecasting because our tests were flawed. Fidel-TS is a new, honest ruler. It shows that:

  1. We need to stop using old, contaminated data.
  2. We need to stop giving models "cheat sheets" (future text).
  3. Current "smart" AI models aren't as good at forecasting as we thought, and we need to build better, specialized tools for the job.

In short, the paper is a call to clean up the kitchen so we can finally see who the real forecasting experts are.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →