OpenSeisML: Open Large-Scale Real Seismic and well-log Dataset for Generative AI
This paper introduces OpenSeisML, a curated collection of real seismic and well-log datasets from the UK National Data Repository, featuring an automated pipeline for time-to-depth conversion, to address the scarcity of public data and enable the training of generative AI models for seismic inversion and uncertainty quantification.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot chef how to bake the perfect cake. To do this, the robot needs to taste thousands of real cakes to understand how flour, sugar, and eggs interact. However, in the world of oil and gas exploration, the "cakes" are underground rock formations, and the "recipes" are hidden inside private company vaults. Most of the high-quality data needed to train these AI chefs is locked away, leaving researchers to practice on fake, computer-generated cakes that don't quite taste like the real thing.
This paper introduces OpenSeisML, a new, open-source "kitchen" filled with real, high-quality seismic data and well logs, designed specifically to train the next generation of AI for finding oil and gas.
Here is a breakdown of what they did, using simple analogies:
The Problem: The "Fake Cake" Dilemma
For a long time, researchers have tried to train AI to figure out what's underground using synthetic data (computer simulations).
- The Compass Model & SEAM: These are like very detailed, realistic-looking fake cakes. They look good, but there's only one version of each. If you train your AI on just one fake cake, it memorizes that specific cake and fails when it sees a slightly different real one.
- OpenFWI: This is like a massive library of fake cakes with different shapes. While there are many of them, they are still made in a controlled lab. They lack the messy, chaotic "flavor" of real geology.
The result? AI models trained on these fakes often struggle when they encounter the messy reality of the actual Earth.
The Solution: The "Real Cake" Library
The authors went to the UK National Data Repository (NDR), a government-run public library of real-world geological data. They built an automated pipeline to grab real seismic surveys (3D images of the underground) and real well logs (measurements taken from actual drill holes).
Think of this pipeline as a super-efficient assembly line that cleans and prepares these raw ingredients so an AI can eat them immediately.
The Assembly Line: How They Prepared the Data
The paper describes a four-step process to turn raw data into training material:
Gathering the Ingredients (Data Collection):
They downloaded massive 3D seismic datasets and associated well logs from the UK. They were picky, mostly choosing data from after the 1990s because older data was like "expired ingredients"—messy, inconsistent, and hard to use automatically.Aligning the Map (Coordinate System):
Imagine trying to overlay a satellite photo of a city with a hand-drawn map of a subway system. If the scales don't match, the lines won't line up. The researchers made sure all the seismic data and the well logs were on the exact same coordinate map so the AI knows exactly where the drill hole is relative to the seismic image.The Time-Travel Conversion (Time-to-Depth):
This is the trickiest part. Seismic data often comes in "time" (how long it took a sound wave to bounce back), but drill holes measure "depth" (how many meters down).- The Analogy: It's like trying to convert a recipe that says "cook for 10 minutes" into "cook until the center reaches 200 degrees." You need a conversion chart.
- The Method: They used "checkshot" data (measurements that tell you exactly how long sound takes to travel to specific depths) to build a smooth velocity model. Think of this as creating a "speed map" of the underground. They used a mathematical tool called Radial Basis Functions (RBF) to fill in the gaps between the drill holes, creating a smooth, 3D speed map that acts as a translator between time and depth.
Slicing the Loaf (Creating Training Data):
Once the 3D data was converted to depth, they sliced it into 2D strips that passed right through the drill holes. They then resized these slices to a standard size (like resizing all photos to 1080p) so the AI could learn from them consistently.
The Result: A New AI Chef
The researchers tested this new dataset by training a Diffusion Model (a type of Generative AI that learns patterns to create new images).
- The Test: They trained the AI on just 40 real wells.
- The Outcome: The AI successfully learned the statistical "flavor" of the underground. It could generate new, realistic velocity models that looked very similar to the real ground truth.
- The Goal: The ultimate aim is to use this AI to create thousands of "what-if" scenarios. This helps geologists understand the uncertainty of what lies underground, rather than just giving them one single guess.
Summary
In short, the paper says: "Stop training your AI on fake, single-version simulations. We have built an automated pipeline to harvest, clean, and organize real seismic data from the UK. This allows AI to learn the true, messy, statistical nature of the Earth, leading to better predictions for seismic inversion."
They are currently working with 2D slices and plan to scale this up to 3D and include more wells (up to 1,000) to make the AI even smarter.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.