biomes: An R package for reproducible occurrence-to-biome classification using 31 global biome schemes
The paper introduces *biomes*, an R package that provides harmonized global raster layers for 31 distinct biome schemes and a reproducible, data-driven workflow to help researchers select the most suitable scheme and classify species occurrence records for ecological and evolutionary studies.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
On a map of the world, the land is often divided into broad, sweeping zones of vegetation. These zones, known as biomes, are the great ecological neighborhoods of our planet, ranging from the dense, humid rainforests of the tropics to the dry, open savannas and the frozen expanses of the polar regions. Scientists use these categories to understand how life is distributed across the globe, how species adapt to their environments, and how ecosystems might change in the future. However, there is a significant problem with how these zones are currently defined. Just as different cartographers might draw the borders of a country in slightly different ways, ecologists have created dozens of different maps for these biomes. Some maps focus on the climate that drives plant growth, others on the physical shape of the vegetation, and still others on how humans have altered the landscape. Because these maps often disagree on where one biome ends and another begins, researchers frequently struggle to choose the right map for their specific question. This uncertainty can lead to confusion, where a study might accidentally use a map that distorts the very patterns it is trying to measure, or where different scientists studying the same group of animals or plants arrive at contradictory conclusions simply because they used different definitions of the world around them.
To solve this confusion, a team of researchers has developed a new digital tool called the biomes package, designed to bring order to this chaotic landscape of maps. The core of their work is a collection of thirty-one distinct global biome schemes, all standardized and prepared for use within the R programming language, a standard tool for ecological data analysis. Instead of forcing scientists to manually download, clean, and align these different maps by hand, the package provides them in a single, harmonized format. More importantly, the package does not just offer the maps; it offers a method to decide which one is best for a specific study. The researchers created a workflow that takes a list of where a species has been found and tests it against all available maps. It then ranks these maps based on how well they fit the data, looking for the scheme that captures the most records while still distinguishing between the different types of environments the species actually occupies. This approach allows researchers to move away from blindly picking a familiar map and instead choose the one that is mathematically and conceptually best suited to their specific research question.
The team demonstrated the power of this new system by applying it to a group of plants known as Bombacoideae, a subfamily that includes the famous baobab tree and the kapok tree. They gathered over 17,000 recorded locations for 185 different species within this group. When they ran their data through the biomes workflow, the system evaluated the thirty-one available maps and, after filtering for those that focused on vegetation types, identified a specific historical map by Ramankutty and Foley as the best fit. This map, which reconstructs what the natural vegetation would look like without human interference, successfully sorted the plant records into meaningful categories. The analysis revealed that nearly half of all the recorded plants, and the vast majority of the species, were found in tropical evergreen woodlands. Another significant portion was found in savannas. These findings matched what scientists already suspected from previous studies, confirming that the group is indeed split between forest and savanna environments. By using the biomes package, the researchers were able to reproduce this known pattern quickly and transparently, without needing to write complex, custom code or manually curate the data.
The value of this work lies in its ability to make the choice of a biome map a deliberate, data-driven decision rather than a default habit. Before this tool, a researcher might have picked a map simply because it was the first one they found or the one everyone else used, potentially introducing errors if that map did not align with their specific study system. The new package ensures that the map chosen is the one that covers the most of the study's data points while still providing enough detail to see the differences between environments. It handles the messy reality of scientific data, such as records that fall on coastlines or islands that some maps ignore, by assigning them a clear status rather than losing them. The system is fast, processing thousands of records in seconds, and it produces clear visualizations that show exactly where the plants are located in relation to the chosen biome boundaries. By providing a standardized way to access and compare these thirty-one different ways of seeing the world's vegetation, the biomes package helps ensure that ecological research is more reproducible, more accurate, and better suited to the specific questions scientists are trying to answer.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.