← Latest papers
💻 computer science

TCBench: A Benchmark for Tropical Cyclone Track and Intensity Forecasting at the Global Scale

TCBench is a global-scale benchmark that enables fair, model-agnostic evaluation of tropical cyclone track and intensity forecasts by leveraging the IBTrACS dataset to compare state-of-the-art physics-based and AI weather prediction models, ultimately aiming to democratize data-driven forecasting and improve extreme event prediction.

Original authors: Milton Gomez, Marie McGraw, Saranya Ganesh S., Frederick Iat-Hin Tam, Ilia Azizi, Samuel Darmon, Monika Feldmann, Stella Bourdin, Louis Poulain--Auzéau, Suzana J. Camargo, Jonathan Lin, Dan Chavas, Ch
Published 2026-06-11
📖 4 min read☕ Coffee break read

Original authors: Milton Gomez, Marie McGraw, Saranya Ganesh S., Frederick Iat-Hin Tam, Ilia Azizi, Samuel Darmon, Monika Feldmann, Stella Bourdin, Louis Poulain--Auzéau, Suzana J. Camargo, Jonathan Lin, Dan Chavas, Chia-Ying Lee, Ritwik Gupta, Andrea Jenney, Tom Beucler

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine trying to predict the path and strength of a hurricane. For decades, scientists have used massive, complex computer simulations based on the laws of physics (like fluid dynamics and thermodynamics) to do this. These are like super-accurate, but very heavy and expensive, "physics engines" that require giant supercomputers to run.

Recently, a new type of "AI engine" has emerged. These are neural networks that learn from past weather data to predict the future, much like how a chess player learns from thousands of past games rather than calculating every possible move from first principles. These AI models are incredibly fast and cheap to run.

The Problem:
While these AI models are great at predicting general weather, they have struggled to accurately predict the specific path and intensity of tropical cyclones (hurricanes/typhoons). Why?

  1. Different Rules: Everyone was playing by different rules. Some researchers used one method to find the storm in the data, others used a different method. It was like comparing apples to oranges.
  2. The "Intensity" Gap: AI models were good at saying where the storm would go, but bad at saying how strong it would get. They often underestimated the wind speeds or missed the "rapid intensification" (when a storm suddenly gets much stronger very quickly).
  3. Data Mismatch: The AI models were often trained on "smoothed out" data that didn't capture the violent, chaotic details of a real hurricane.

The Solution: TCBench
The authors of this paper created TCBench, which acts like a universal referee and a standardized testing ground.

  • The "Rulebook": They took all the messy, different data sources (observations, physics models, and AI models) and forced them into a single, standardized format. They used a "ground truth" dataset called IBTrACS (the official record of where storms actually went) as the scorecard.
  • The "Race": They set up a fair race using data from 2023. They took the best physics-based models and the best new AI models and asked them the same question: "Given a storm exists right now, where will it be in 1, 2, 3, 4, or 5 days, and how strong will it be?"
  • The "Spotter": Since AI models sometimes "lose track" of a storm (they might predict the storm disappears or never existed), the benchmark has a safety net. If an AI model fails to predict a storm, it defaults to a simple "persistence" guess (assuming the storm just keeps doing what it's doing right now) so the comparison remains fair.

What They Found (The Results):

  1. The Path (Track): The AI models are excellent at predicting where the storm will go. In fact, for predicting the path, the AI models are just as good as, and sometimes even better than, the traditional physics supercomputers. They are like a GPS that knows the road better than anyone else.
  2. The Strength (Intensity): The AI models were initially bad at predicting how strong the storm would get. They often guessed the winds would be weaker than they actually were.
    • The Fix: However, the authors showed that if you take the AI's raw guess and run it through a simple "post-processing" step (like a coach giving the player a quick pep talk and a few tactical adjustments), the AI's intensity predictions become very accurate.
  3. The "Sudden Explosion" (Rapid Intensification): Predicting when a storm will suddenly get much stronger is the hardest challenge. The raw AI models mostly missed this. Only the AI models that were specifically tweaked or "post-processed" could spot these dangerous spikes in strength.

The Big Picture:
TCBench isn't just a list of scores; it's a toolkit. It provides the data, the rules, and the measuring sticks so that scientists and AI developers can stop arguing about who is right and start working together to make better predictions.

The paper concludes that while AI models are already winning the race for predicting where storms go, they need a little help (like post-processing or specific training) to get good at predicting how strong they will get. By lowering the barrier to entry and standardizing the tests, TCBench hopes to speed up the development of tools that can save lives by giving us better warnings about these destructive storms.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →