← Latest papers
🤖 machine learning

Towards Reliable Zero-Shot Crowd Forecasting: Evaluating Time Series Foundation Models for Special Event Pedestrian Forecasting

This paper evaluates the effectiveness of pretrained time series foundation models for zero-shot probabilistic crowd forecasting during infrequent special events, demonstrating their ability to provide reliable point estimates and uncertainty quantification without extensive local retraining.

Original authors: Ziteng Li, Yanan Xin, Tina Comes, Serge Hoogendoorn

Published 2026-07-21
📖 5 min read🧠 Deep dive

Original authors: Ziteng Li, Yanan Xin, Tina Comes, Serge Hoogendoorn

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to predict the weather for a picnic that only happens once every five years. You can't just look at last year's weather because, well, it didn't happen last year. You also can't wait until the day of to start learning the patterns. This is the tricky world of "zero-shot" forecasting: making smart guesses about the future when you have almost no past data to study. Usually, computers need to eat mountains of historical data to learn how to predict things like traffic or crowds. But what if you could use a super-smart computer brain that has already read the "encyclopedia" of time and patterns from millions of other places, and just ask it, "Hey, what's going to happen here?" This is the promise of "foundation models." They are like universal students who have already done their homework on a billion different topics, so they can walk into a new classroom and start teaching immediately without needing a textbook.

Now, picture a massive maritime festival in Amsterdam, the SAIL event, which draws 2.5 million people. The city needs to know exactly where the crowds will be in the next 30 minutes to open extra gates, send more police, or clear a path for an ambulance. But because this event is rare, the city doesn't have years of data to train a custom robot. They need a forecast that is not just a single guess (like "it will be 500 people"), but a range of possibilities with a safety net (like "it will be between 400 and 600, but be careful, it might spike to 800"). This paper is a real-world test drive to see if these pre-trained "universal students" can actually handle the chaos of a massive, one-off crowd event without crashing or giving dangerous advice.

The researchers set up a race between two of the smartest time-traveling AI models available: TimesFM and Chronos-2. They didn't train these models on Amsterdam data; instead, they dropped them straight into the middle of the SAIL2025 festival to see how they performed in a "zero-shot" setting. Think of it like handing a chef who has never cooked in a specific kitchen a new, strange recipe and seeing if they can make a delicious meal without tasting the ingredients first.

The results were surprisingly clear. The Chronos-2 model turned out to be the star of the show. It managed to predict crowd flows with a reliable "safety window" of about 30 to 45 minutes ahead. In the world of crowd control, having a 30-minute heads-up is like having a superpower; it gives managers enough time to actually move people around or open emergency exits before a bottleneck becomes a disaster. Chronos-2 was particularly good at "rolling" its predictions forward, meaning as new data came in every 90 seconds, it updated its forecast smoothly without getting confused.

However, the paper also found some interesting quirks. While Chronos-2 was the overall champion, TimesFM actually did a better job during the quiet, low-crowd hours (like late at night) and in areas where the crowd numbers were very small and steady. It's as if TimesFM is a calm, steady librarian who is great at predicting quiet library hours, while Chronos-2 is a high-energy sports coach who excels when the game gets chaotic and fast-paced.

One of the most critical things the paper discovered is how these models handle "missing data." Imagine if a camera stopped working for an hour, leaving a blank spot in the crowd count. The researchers found that Chronos-2 reacted to this missing information by widening its "safety net." Instead of confidently guessing a specific number, it said, "I don't know exactly, so the range of possibilities is now huge." This is a feature, not a bug! In safety-critical situations, it is far better for a model to say "I'm unsure, be careful" than to confidently give a wrong answer. TimesFM, on the other hand, didn't widen its net as much, which could be riskier if the missing data hid a sudden surge.

The study also looked at the "cost" of using these models. They ran everything on a standard laptop without fancy graphics cards. The good news? Both models were incredibly fast, taking less than half a second to predict an hour into the future. This means they could easily run on regular city computers to help manage crowds in real-time. The memory they used was also manageable, though Chronos-2 needed a bit more power to "wake up" and start thinking.

In the end, the paper suggests that for managing rare, massive events like the SAIL festival, Chronos-2 is the more reliable choice, especially when you need to handle sudden changes and missing data. It suggests that we don't always need to build a new, custom robot for every single event; sometimes, we just need to find the right pre-trained expert who can jump in and handle the job. The researchers also introduced a new set of "report cards" for these models, focusing not just on how close the guess was, but on how stable the prediction was over time and how honest the model was about its uncertainty. This is a big deal because, in crowd management, knowing when you don't know is just as important as knowing the answer.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →