Earthquake aftershock forecasting with conditional generative models
The paper introduces QuakeGen, a conditional diffusion model that outperforms traditional statistical methods by generating spatiotemporal aftershock fields that accurately capture fault-controlled anisotropic patterns and variable productivity, thereby advancing earthquake forecasting through data-driven generative modeling.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
When a massive earthquake strikes, the ground does not simply go quiet. For hours, days, and sometimes years afterward, the crust continues to shudder, releasing thousands of smaller tremors known as aftershocks. These follow a predictable rhythm: they happen most frequently right after the main event and gradually fade away, while their locations tend to cluster along the specific fault line that broke. Scientists have long used statistical rules to guess where and when these aftershocks might occur, treating each tremor as an isolated point in a vast, random cloud. These traditional methods work well on average, but they often fail to capture the messy, real-world details of how a specific sequence unfolds. They struggle to predict the exact shape of the aftershock zone or how many tremors a specific sequence will produce, often smoothing over the jagged, fault-aligned patterns that define a real earthquake's aftermath.
A researcher at the University of California, Berkeley, has developed a new way to forecast these sequences that moves beyond simple statistics. Instead of trying to predict individual tremors one by one, they trained a computer model to imagine the entire future landscape of seismic activity as a single, evolving picture. This new approach, called QuakeGen, learns from vast amounts of historical earthquake data to understand how aftershocks spread across space and time. When tested against the standard methods used by government agencies, this new model proved more accurate at pinpointing where aftershocks would cluster and how the activity would decay. It successfully recreated the long, narrow shapes of aftershock zones that follow real fault lines, a detail that older models often miss, replacing them with vague, circular guesses.
The core idea behind this work is a shift in perspective. Traditional forecasting treats an earthquake sequence like a list of separate events, using fixed mathematical formulas to estimate how many will happen and where. These formulas assume that aftershocks spread out evenly in all directions from the main shock, like ripples in a pond. However, real earthquakes do not behave this way. The aftershocks usually line up along the broken fault, creating long, stretched-out clusters that reflect the geometry of the rupture deep underground. The researcher realized that trying to predict every single event individually was limiting their ability to see this bigger picture. Instead, they decided to treat the forecast as a field, similar to how meteorologists map the weather. They wanted a model that could generate a complete map of future seismic activity, showing both the number of expected tremors and the size of the largest one in every small patch of ground.
To build this system, the researcher used a type of artificial intelligence known as a diffusion model. This technology, which has recently revolutionized fields like weather prediction and protein structure analysis, works by learning to reverse a process of adding noise. Imagine a clear photograph that is slowly covered in static until it becomes pure, meaningless grain. A diffusion model learns how to take that grain and, step by step, remove the noise to reveal the original image. In this study, the "image" is not a photograph but a map of earthquake activity. The model starts with a blank, noisy map and, guided by the information of what has already happened, gradually refines it into a clear forecast of what will happen next.
The researcher trained their model, QuakeGen, on two different types of data to see how well it could generalize. First, they used a global dataset containing thousands of mainshocks and their aftershocks from around the world, spanning from 1990 to 2023. They taught the model to look at the seismic activity in the days leading up to a large earthquake and the first few hours of aftershocks immediately after. Based on this limited information, the model had to predict the pattern of aftershocks for the next several days, weeks, or even months. They then tested the model on 80 major earthquakes that occurred in 2024 and 2025, which the model had never seen before. In these tests, QuakeGen outperformed the standard operational model used by the United States Geological Survey. While the traditional model produced a broad, circular halo of predicted activity, QuakeGen correctly identified the elongated, fault-aligned shapes of the aftershock zones. For example, after a magnitude 7.5 earthquake in Japan, the new model reproduced the 150-kilometer-long stretch of aftershocks along the fault, whereas the old model spread the prediction out in a wide circle.
The second test focused on a single region: southern California. Here, the goal was not just to predict aftershocks of a single big quake, but to forecast the daily rate of earthquakes across a large area over time. The researcher trained the model on a highly detailed catalog of seismicity in the region, which includes many small tremors that standard detectors often miss. They then asked the model to predict the daily earthquake counts for specific sub-regions known for frequent activity, such as the Salton Sea and the San Jacinto fault. In this setting, the model matched the performance of the best-tuned statistical models that have been carefully adjusted for that specific area. This is significant because it suggests that a single, data-driven model can learn the complex behavior of an entire fault system without needing to be manually re-calibrated for every new region.
One of the most striking aspects of the new model is how it handles uncertainty. Because the future of an earthquake sequence is inherently unpredictable, the model does not produce just one single answer. Instead, it generates an ensemble of 100 different possible futures for each forecast. Each of these 100 maps is a valid realization of what could happen, showing different specific locations for individual tremors. When the researcher averaged these 100 maps together, they got a smooth picture of the overall rate of activity, which closely matched the actual patterns observed in nature. This approach allows forecasters to see not just the most likely outcome, but the full range of possibilities, providing a clearer picture of where the risk is concentrated.
The study also explored the model's ability to predict the size of the largest expected aftershock in any given area. Traditional methods often struggle with this, as they rely on statistical averages that can be skewed by empty areas. QuakeGen, however, generates the count of tremors and the size of the largest one simultaneously. By looking at the 90th percentile of its predictions, the model was able to concentrate the forecast for the largest tremors along the ruptured fault, matching the real-world observation that the biggest aftershocks tend to happen right where the main break occurred. This joint prediction of frequency and magnitude offers a more complete and realistic view of the hazard.
Despite these successes, the researcher acknowledges that the model has limits. It relies heavily on the quality and quantity of the data it is trained on. For the very largest earthquakes, which are rare, the model sometimes underestimates the total number of aftershocks or fails to capture the full extent of the rupture zone. This is because the model has seen fewer examples of these massive events during its training. The author suggests that future improvements could come from combining this data-driven approach with physics-based simulations to generate more examples of great earthquakes, or by feeding the model additional information, such as the exact geometry of the fault rupture or measurements of ground deformation.
The work represents a significant step forward in how scientists approach earthquake forecasting. By moving away from fixed rules and isolated event predictions, and toward learning the complex, evolving patterns of seismic fields directly from data, the researcher has created a tool that captures the true nature of earthquake sequences. The model does not just guess; it learns the underlying structure of how faults behave. While it cannot predict when the next big earthquake will strike, it offers a much sharper, more accurate view of what happens after one does, helping communities understand the shape and scale of the shaking that follows.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.