← Latest papers
🔬 optics

Beyond chlorophyll: machine learning estimates of diagnostic phytoplankton pigments from multispectral ocean colour data

This study demonstrates that machine learning models trained on multispectral satellite ocean-colour data can more accurately estimate diagnostic phytoplankton pigment concentrations than traditional chlorophyll-a-only approaches, thereby improving the large-scale observation of phytoplankton community composition.

Original authors: David Moffat, Angus Laurenson, Victor Martinez-Vicente, Gemma Kulk, Xuerong Sun, Robert J. W. Brewin, Shubha Sathyendranath

Published 2026-08-25
📖 6 min read🧠 Deep dive

Original authors: David Moffat, Angus Laurenson, Victor Martinez-Vicente, Gemma Kulk, Xuerong Sun, Robert J. W. Brewin, Shubha Sathyendranath

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The ocean is a living engine, driven by microscopic plants called phytoplankton that float in sunlit waters. These tiny organisms are the foundation of the marine food web and are responsible for roughly half of all the organic carbon produced on Earth through photosynthesis. To understand how these plants function and how they might respond to a changing climate, scientists need to know not just how much of them are present, but what kinds they are. For decades, the standard way to measure their abundance from space has been to look for a specific green pigment called chlorophyll-a. This pigment is found in almost all phytoplankton and acts as a reliable proxy for total biomass, much like counting the leaves on a tree to estimate its size. However, knowing the total amount of leaves does not tell you if the tree is an oak, a pine, or a willow. In the ocean, different groups of phytoplankton play different roles in the ecosystem, and distinguishing between them is crucial for understanding the health of the planet.

The challenge lies in the fact that the ocean is a complex mixture of light and life. While chlorophyll-a is easy to spot from space, other pigments that act as unique fingerprints for specific groups of phytoplankton are much harder to detect. These accessory pigments help the plants adapt to different light conditions, but their signals are often drowned out by the overwhelming presence of chlorophyll-a or are difficult to separate using the limited number of color bands that current satellites can see. For years, researchers have wondered if the subtle differences in the colors of the ocean, beyond just the intensity of the green, hold enough information to reveal the identity of the microscopic communities living within. A new study by David Moffat and his colleagues at the Plymouth Marine Laboratory and the University of Exeter sets out to answer this question by using advanced computer learning techniques to look deeper into the data.

The researchers began with a massive collection of real-world measurements. They gathered over 33,000 samples of seawater taken from ships and research vessels around the globe between 1991 and 2021. In a laboratory, these samples were analyzed to identify the exact concentrations of various pigments, creating a detailed ground truth of what was actually in the water. They then matched each of these samples with satellite images taken on the same day and in the same location. The satellite data came from the European Space Agency's Ocean Colour Climate Change Initiative, which provides a consistent record of ocean colors from multiple sensors over many years. The team used this matched dataset to train two different types of machine learning models. One model was a sophisticated foundation model designed to learn from tabular data, while the other was a random forest model, a type of algorithm known for its ability to find complex patterns without overfitting. To test if these models were truly learning something new, they also built a simple baseline model that relied only on the concentration of chlorophyll-a, ignoring all other color information.

The results showed that the models using the full spectrum of ocean colors consistently outperformed the simple chlorophyll-only approach. This finding confirms that the ocean's color contains more information than just the amount of plant life; it holds clues about the specific types of plants present. However, the success of the models varied depending on which pigment they were trying to identify. For fucoxanthin, a pigment often associated with diatoms, the models performed very well, but this was largely because fucoxanthin is so strongly linked to the total amount of chlorophyll-a that the simple baseline model could already predict it with reasonable accuracy. The real breakthrough came with other pigments, such as alloxanthin and certain forms of fucoxanthin that are markers for different groups. For these, the models that used the full range of color data showed a dramatic improvement over the baseline. The computer learned to spot subtle shifts in the ocean's color that were invisible to the simple chlorophyll count, allowing it to distinguish between different phytoplankton communities with much greater precision.

To understand how the computer was making these decisions, the researchers used a technique that breaks down the model's logic to see which pieces of information mattered most. They found that for some pigments, the model relied heavily on the total amount of chlorophyll-a, confirming that these pigments move in lockstep with the overall biomass. But for others, the model ignored the total amount of green and instead focused on the specific ratios of blue, green, and red light reflected by the water. This suggests that the models were not just guessing based on how much plant life was there, but were actually detecting the unique optical signatures of specific groups. When the researchers applied these trained models to create global maps of pigment distribution, the results painted a picture that matched known ecological patterns. The maps showed high concentrations of certain pigments in nutrient-rich coastal waters and upwelling zones, while other pigments dominated the vast, nutrient-poor tropical oceans. This spatial separation mirrors the known habits of different phytoplankton groups, with some thriving in rich, turbulent waters and others adapted to the calm, clear conditions of the open ocean.

The study also highlighted the limitations of current technology. While the models could successfully separate many groups, they struggled with zeaxanthin, a pigment found in very small, abundant bacteria-like organisms. The models could not predict the exact amount of this pigment very well, likely because its signal is weak and difficult to separate from the background noise in the satellite data. Nevertheless, even for this difficult pigment, the model managed to capture a broad pattern that matched the expected distribution of these tiny organisms in warm, tropical waters. This indicates that while the technology is not yet perfect, it is already capable of extracting meaningful ecological information that goes far beyond simple biomass counts. The researchers emphasize that these findings are not a final solution but a significant step forward. They demonstrate that machine learning can unlock hidden details in decades of satellite records, offering a new way to monitor the changing composition of the ocean's microscopic life. As new satellites with more sensitive color sensors are launched, these methods will become even more powerful, providing a clearer window into the complex and vital world of phytoplankton that sustains our planet.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →