A harmonised dataset for Earth system foundation models
This paper introduces WorldTensor, a harmonised global dataset that aligns diverse environmental and socioeconomic variables onto a standardised 0.25° grid and annual framework to enable the training of multimodal Earth system foundation models that capture coupled human-environment dynamics.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the Earth as a giant, complex machine with two main engines: the natural world (weather, oceans, forests) and the human world (cities, farms, power plants, economies). For a long time, scientists have built "super-brains" (called foundation models) to understand how these engines work. However, there's been a major problem: the super-brains have only been fed data about the natural engine. They know everything about rain and wind, but they are almost blind to how humans build cities, burn fuel, or move money.
This paper introduces WorldTensor, a massive new "instruction manual" designed to fix this blind spot. Think of WorldTensor not as a single book, but as a giant, perfectly organized library where every single fact about the planet—from a drop of rain to a factory smokestack—is written on the same type of paper, in the same language, and filed in the same drawer.
Here is how the paper breaks it down:
1. The Problem: A Messy Pile of Puzzles
Before WorldTensor, trying to study the Earth and humans together was like trying to solve a puzzle where half the pieces are from a picture of the ocean, and the other half are from a picture of a city, but they are all different sizes, different shapes, and printed in different languages.
- Some data is daily (like weather); some is only every few years (like population counts).
- Some data is a grid of squares (like a map); some is just a list of dots (like power plant locations).
- Some maps use different coordinate systems, making it impossible to line them up.
Scientists had to spend months just cleaning and aligning this data before they could even start their research.
2. The Solution: The "WorldTensor" Library
The authors created WorldTensor, a harmonized dataset that acts like a universal translator and a master organizer.
- The Grid: They took hundreds of different data sources and forced them all onto a single, standard grid (0.25 degrees, roughly the size of a large city block). Imagine taking a high-resolution photo, a sketch, and a list of numbers, and printing them all on the exact same graph paper so they line up perfectly.
- The Time: They standardized time to a yearly rhythm. If a source had daily data, they summarized it into a yearly average. If a source only had data every 5 years, they filled in the gaps to make a smooth yearly timeline.
- The Content: It's a "multimodal" mix. It doesn't just have weather; it has:
- Nature: Oceans, ice, soil, forests, and air quality.
- Humans: Population, GDP (wealth), power plants, roads, and even conflict zones.
- The Mix: How humans change nature (like cutting down forests) and how nature affects humans (like floods hitting cities).
3. How They Built It (The Kitchen Analogy)
Think of the data sources as raw ingredients from different markets. Some are fresh fish (daily weather), some are dried beans (annual census), and some are spices (point locations of factories).
- Regridding: They took the "fish" and chopped it into the same size pieces as the "beans."
- Rasterizing: They took the "spices" (dots on a map) and spread them out so they looked like a smooth layer of seasoning on the grid.
- Standardizing: They made sure every ingredient was measured in the same units (like converting cups to grams) and labeled clearly.
- Quality Control: They checked the ingredients to make sure they weren't spoiled. If a number was physically impossible (like a temperature of 500°C on Earth), they flagged it. They also checked that the "map" didn't have tears or seams where the edges didn't match.
4. What It Looks Like
The paper shows that this dataset is huge. It contains over 52,000 files covering the years 1900 to 2025.
- It includes 14 different "departments" (like Climate, Energy, Human Systems, Hazards).
- It allows a computer to see that a change in emissions (human) is linked to a change in air quality (nature), or that urbanization (human) is linked to flood risk (nature).
5. Why It Matters (The "Super-Brain" Training)
The main goal of WorldTensor is to train Foundation Models.
- Before: A super-brain learned to predict the weather but didn't understand that building a dam changes the river flow, or that a war stops people from farming.
- Now: With WorldTensor, a super-brain can learn the coupled dynamics. It can learn how human actions and natural forces push and pull against each other on a global scale.
The paper validates this by showing that the data works. When they tested it against real historical events (like the 1997 El Niño or the 2020 pandemic), the data correctly showed the spikes and dips in temperature, emissions, and conflict. It proved that the "library" is accurate and ready for computers to read.
Summary
WorldTensor is a massive, clean, and organized dataset that finally puts the "Human" and "Earth" engines on the same page. It turns a chaotic pile of mismatched maps and numbers into a single, coherent picture of our planet, allowing AI to finally learn how we live on and with the Earth, rather than just studying the Earth in isolation.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.