← Latest papers
🤖 machine learning

Earth Observation Foundation Models for Terrestrial Ecohydrology: From Representation Learning to Process Inference

This paper proposes a process-aware framework to address the misalignment between Earth observation foundation models and ecohydrological needs by highlighting current gaps in data coverage, validation depth, and physical consistency while advocating for targeted evaluation to enable trustworthy monitoring of coupled water, energy, and carbon dynamics.

Original authors: Yi Yu, Jian Peng, Yucheng Lin, Trevor F. Keenan, Thomas F. A. Bishop

Published 2026-08-18
📖 6 min read🧠 Deep dive

Original authors: Yi Yu, Jian Peng, Yucheng Lin, Trevor F. Keenan, Thomas F. A. Bishop

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Earth is a living system where water, energy, and carbon move constantly between the soil, the plants, and the air. When rain falls, it might soak into the ground, run off into a river, or evaporate back into the sky. Plants drink from the soil and release water vapor through their leaves, a process that cools the air and pulls carbon dioxide from the atmosphere. Understanding these invisible exchanges is vital for predicting droughts, managing water supplies, and figuring out how ecosystems will respond to a changing climate. For decades, scientists have used satellites to watch this dance from space, capturing images of the land in different colors and wavelengths to measure things like soil wetness, plant health, and surface temperature. However, turning these raw satellite pictures into reliable answers about how the Earth works has always been difficult, because the data is messy, the patterns are complex, and the conditions on the ground change constantly.

Recently, a new type of artificial intelligence called a foundation model has emerged in the field of Earth observation. Think of these models as massive, pre-trained systems that have studied billions of satellite images to learn general patterns about the planet, much like a student who has read every book in a library before taking a specific exam. The hope has been that these models could be quickly adapted to solve specific problems, such as predicting how much water a forest is using or how dry the soil is in a particular region. But a team of researchers from the University of Sydney and other institutions asked a critical question: do these powerful new tools actually work for the specific, complex science of ecohydrology, or are they just good at simpler tasks? They set out to investigate whether these models can truly help us understand the deep connections between water, energy, and life on land, or if they are falling short in ways that matter.

The researchers began by mapping out exactly how scientists usually turn a satellite signal into a scientific fact. They found that this process happens in steps, moving from simple measurements to complex conclusions. At the bottom level, a satellite records raw light or heat. The next step involves turning that light into a basic measurement, like the temperature of the ground or the amount of green vegetation. The third step is where things get harder: scientists combine those basic measurements with weather data and computer models to estimate hidden states, such as how much water is deep in the soil or how much carbon a forest is absorbing. The final step is using all that information to make decisions about droughts or water management. The team realized that for a foundation model to be useful, it must be able to handle the uncertainty and complexity at every single one of these steps, not just the easiest ones.

To see if current models meet these demands, the team conducted a massive review of the scientific literature, analyzing dozens of foundation models released between 2021 and 2026. They discovered a significant gap between what these models are trained on and what ecohydrologists need to know. Most of these models are trained primarily on visible light images and active radar data, which are good for seeing shapes and structures. However, they are largely missing data from thermal sensors, which measure heat, and passive microwave sensors, which are crucial for seeing through clouds to measure soil moisture. Even more importantly, none of the models they reviewed were trained on data that directly measures the fluorescence of plants, a subtle signal that reveals how well plants are photosynthesizing. This means the models are learning from an incomplete picture of the Earth's water and energy cycles.

When the researchers looked at how these models were actually being used in real-world studies, the results were mixed. The models showed great promise when the task was to provide general context, such as helping to identify different types of land cover or improving the spatial detail of a map. They were also very good at helping scientists work with very little data, allowing them to make predictions where few measurements existed. However, as the tasks became more complex and required deep understanding of physical processes, the models struggled. For example, when researchers tried to use these models to predict the exact amount of water flowing in a river or the precise rate at which plants are exchanging carbon with the air, the results were often weak or unreliable. The models tended to perform well only when the conditions were similar to the data they had seen during training, and they often failed when asked to predict events in new regions or under extreme weather conditions.

The study also highlighted a problem with how these models are tested. Many of the existing tests rely on comparing the model's output to other computer-generated maps rather than checking against real measurements taken on the ground. This creates a cycle where the models are just learning to mimic other imperfect models rather than discovering the truth. The researchers found that very few studies tested the models against independent, high-quality data from weather stations or field sensors. Furthermore, the tests rarely checked if the models respected the basic laws of physics, such as the rule that water cannot be created or destroyed. Without these rigorous checks, it is difficult to trust the models when they are used to make important decisions about water resources or climate adaptation.

The authors conclude that while these foundation models are powerful tools, they are not yet ready to replace the careful, physics-based methods that scientists have used for decades. The models are excellent at providing a broad, spatial context and can help fill in gaps where data is missing, but they cannot yet reliably infer the deep, causal mechanisms that drive the water and carbon cycles. To move forward, the researchers argue that the development of these models must change. Future models need to be trained on a wider variety of data, including heat and passive microwave signals, and they must be tested against real-world measurements, not just other computer models. They also need to be designed with the laws of physics in mind, ensuring that their predictions make sense in the real world. Only by aligning these powerful new tools with the specific needs of ecohydrology can we hope to use them to build a more trustworthy understanding of our planet's changing water and energy systems.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →