← Latest papers
🤖 AI

LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models

This paper introduces LiveHouse-TS, the first open-world living benchmark that evaluates Time Series Foundation Models through continuous prequential testing on real future data, revealing that static rankings often fail to reflect long-term robustness and performance under evolving distribution shifts.

Original authors: Haomin Wen, Ziyu Zhou, Qingxiang Liu, Siru Zhong, Yuxuan Liang

Published 2026-08-19
📖 5 min read🧠 Deep dive

Original authors: Haomin Wen, Ziyu Zhou, Qingxiang Liu, Siru Zhong, Yuxuan Liang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Predicting the future is a fundamental human urge, whether we are checking the weather, planning a budget, or trying to understand how a river will flow after a storm. For decades, scientists have built mathematical models to forecast these patterns, but a new generation of tools has recently emerged that promises to do this without needing to be taught the specifics of every single situation. These tools, known as foundation models, are trained on vast amounts of historical data from many different fields at once. They learn the general language of time and change, allowing them to make educated guesses about new, unseen data immediately, a capability called zero-shot forecasting. The hope has been that these powerful systems could replace the need for custom-built models for every specific task, offering a universal solution for predicting everything from stock prices to energy usage.

However, there is a significant gap between how these models are tested in the lab and how they actually perform in the real world. Traditionally, researchers evaluate these systems using static snapshots: they take a fixed chunk of past data, train a model, and then see how well it predicts a fixed chunk of future data that was set aside. This approach is like grading a student on a single practice test that never changes. While it provides a clear score, it fails to capture the messy reality of life, where conditions shift constantly, seasons change, and unexpected events occur. A model might ace a static test by memorizing patterns that no longer exist, only to fail when the world moves on. The question remains: do these models remain reliable when the data they are predicting is constantly evolving, or do they crumble when the ground rules change?

To answer this, a team of researchers has introduced a new way of testing called LiveHouse-TS. Instead of a frozen snapshot, they built a living, breathing benchmark that operates like a continuous performance review. Imagine a stage where models must make predictions in real-time, at the exact moment the data is available, before the true outcome is known. As new data arrives every day, the models are asked to forecast again, and their performance is scored immediately. This process repeats endlessly, creating a stream of evaluations that mimics the actual conditions of deployment. The system covers seventeen different datasets from eleven distinct domains, ranging from weather and ocean waves to financial markets and earthquake counts. It tests models across various time scales, from seconds to years, ensuring that the evaluation captures a wide diversity of real-world behaviors.

The researchers ran this live experiment with several of the most advanced forecasting models available today, alongside traditional statistical methods. The results revealed a dramatic reshuffling of the rankings. Models that had previously dominated the static, snapshot-based leaderboards found themselves struggling in the live environment. In the static tests, a model might appear to be the best because it perfectly matched a specific historical period. But when subjected to the continuous flow of new data, these same models often degraded quickly, unable to adapt to shifting patterns or unexpected changes. Conversely, some models that were not the top performers in static tests proved to be far more robust, maintaining consistent accuracy as the data evolved.

One of the most striking findings was that a model's ability to predict a single point in time does not guarantee it can handle the uncertainty of the future. The study introduced new ways to measure not just how accurate a prediction is, but how stable it remains over time and whether it improves as more information becomes available. They found that while many modern models could produce excellent point forecasts, they often failed to provide reliable probability estimates or to adjust their confidence levels when the data became volatile. For instance, in domains like weather and ocean dynamics, the top-ranked models from static tests performed poorly, while others adapted better to the changing conditions. This suggests that the current standard of testing, which relies on fixed datasets, may be giving a false sense of security about which models are truly ready for real-world use.

The study also highlighted that no single model is a universal winner. Even the most sophisticated systems showed weaknesses in specific areas, such as handling sudden spikes in data or adapting to long-term trends. The live benchmark revealed that performance is not a permanent attribute; a model that is excellent today might not be tomorrow if the underlying data patterns shift. By continuously updating the leaderboard as new data arrives, the researchers demonstrated that rankings are fluid and must be earned repeatedly. This approach exposes flaws that static tests miss, such as a model's tendency to smooth out important details or its failure to react quickly to new information.

Ultimately, this work suggests that the path to truly reliable forecasting requires moving beyond one-time evaluations. The researchers argue that to understand how these powerful tools will behave in practice, we must test them in the environment where they will actually be used: a world that is constantly changing. The live benchmark provides a framework for this, offering a transparent and reproducible way to see which models can truly stand the test of time. It does not declare a single champion but rather provides a dynamic map of strengths and weaknesses, helping scientists and practitioners choose the right tool for the job based on how it performs under real-world pressure. The findings serve as a reminder that in the complex, shifting landscape of time series data, adaptability is just as important as accuracy.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →