← Latest papers
📊 statistics

Bayesian inference on beta diversity via feature allocation models with imperfect detection

This paper introduces a novel Bayesian feature allocation framework that integrates latent species composition models with imperfect detection mechanisms to provide robust, scalable, and uncertainty-quantified inference on beta diversity and species sharing in high-dimensional ecological data.

Original authors: Federica Stolf, Tommaso Rigon, David B. Dunson

Published 2026-08-12
📖 8 min read🧠 Deep dive

Original authors: Federica Stolf, Tommaso Rigon, David B. Dunson

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery in a giant, invisible city. This city is the natural world, and its citizens are the millions of tiny species of fungi, bacteria, and insects that we can't see with our naked eyes. To understand how this city works, scientists need to know two things: who lives there, and how different the neighborhoods are from one another. This is the heart of a field called ecology, specifically a concept known as "beta diversity." Think of beta diversity as a measure of how much the cast of characters changes as you walk from one park to another. If Park A has only squirrels and Park B has only pigeons, they are very different (high beta diversity). If both parks have squirrels, pigeons, and robins, they are quite similar (low beta diversity).

However, solving this mystery is incredibly tricky. First, the "citizens" are often hiding. Just because you don't see a squirrel in a tree doesn't mean it isn't there; maybe it's just sleeping, or maybe your binoculars aren't good enough. In science, this is called "imperfect detection." Second, the city is huge. There are so many species that even if you look everywhere, you will likely miss the rare ones. Traditional methods for counting these invisible neighbors often act like they have perfect vision, assuming that if they don't see a species, it simply doesn't exist. This leads to a distorted picture of the world, making us think some neighborhoods are empty when they are actually bustling with life.

This is where a new team of statisticians steps in with a fresh set of tools. They have developed a clever mathematical framework called MOSAIC (Multi-site Occupancy-aware Species Allocation with Imperfect deteCtion). Instead of pretending their eyes are perfect, they build a model that admits, "We might have missed some guests." They treat the data like a game of "guess who," where they calculate the probability that a species is actually present but just wasn't spotted, versus the probability that it truly isn't there. By doing this, they can create a much clearer map of how species are shared across different locations, even when the data is messy and full of missing pieces.

The Invisible Guest List

To understand what the authors did, imagine you are hosting a massive, global potluck dinner. You have invited guests from 34 different cities around the world. At each city's table, people are taking photos of the food they brought. But here's the catch: the cameras are glitchy. Sometimes a dish is right there on the table, but the camera fails to snap a picture. Other times, the dish is genuinely missing.

In the past, scientists analyzing these potluck photos would simply count what they saw. If a photo from City A showed a pizza and City B showed a salad, they would say, "These two cities have nothing in common." But this ignores the glitchy cameras. Maybe City B actually had a pizza too, but the camera missed it. This leads to a false conclusion that the cities are more different than they really are.

The MOSAIC model is like a super-smart detective who looks at the photos and asks, "Given how often this camera misses things, how likely is it that a pizza was actually there?" The model separates two distinct events: Occupancy (is the species actually at the site?) and Detection (did we actually see it?). By keeping these two ideas separate, the model can estimate the true number of species present, even the ones that never made it into the photos.

The Magic of "Partial Exchangeability"

One of the biggest hurdles in this kind of detective work is that nature isn't uniform. In many old statistical models, scientists assumed that every location was just a random shuffle of the same pool of species. It was like assuming that every city in the world has the exact same mix of people, just in a different order. But we know that's not true. A city in the tropics has a totally different "vibe" and guest list than a city in the Arctic.

The authors introduce a concept called "partial exchangeability." Think of it like this: instead of assuming every city is a random shuffle of the same deck of cards, they allow each city to have its own unique deck, but with some rules. The decks are related (they all come from the same global pool of species), but the specific cards in each deck depend on the local conditions, like temperature or wind. This allows the model to respect the fact that a tropical forest really does look different from a snowy tundra, while still using information from one place to help guess what's happening in another.

The "Beta-Diversity" Score

The main goal of this paper is to measure beta diversity—how different the guest lists are between two cities. The authors propose a new way to calculate this score that accounts for the glitchy cameras.

They define a score (let's call it the "Similarity Score") that ranges from 0 to 1.

  • 0 means the two cities are identical in their species makeup.
  • 1 means they share absolutely nothing.

Crucially, this score is designed to be fair. If City A has 100 species and City B has 100 species, and they share 50, the score reflects that. But if City A has 1,000 species and City B has 1,000, and they share 50, the score knows that's a much bigger difference. The authors show that their new method is robust: even if you aren't sure exactly how "glitchy" the cameras are (a parameter they call σ\sigma), the final similarity score stays surprisingly stable. It's like having a compass that points North even if you're slightly off-center.

Testing the Theory with Fake Data

Before applying their model to real life, the authors played a game with "fake" data. They created a computer simulation of a world with 3,000 species and 3 different sites. They programmed the cameras to miss things on purpose. Then, they let their MOSAIC model try to guess the true number of shared species.

The results were promising. The model's predictions closely matched the "true" numbers they had programmed into the simulation. It successfully figured out that even though the cameras missed a lot of species, the underlying pattern of who was sharing what was still recoverable. They also tested it on a more complex scenario with 15 sites and 5,000 species, and again, the model held its own, often doing better than older methods that ignored the "glitchy camera" problem.

The Real-World Test: Fungi in the Air

Finally, the authors took their model to the real world using data from the Global Spore Sampling Project (GSSP). This project collected air samples from 34 locations across the Northern Hemisphere, from North America to the Arctic. The goal was to study airborne fungi.

The data was a mess in the best possible way for this kind of problem:

  • They found 15,352 distinct types of fungi.
  • But more than half of them (57%) were "one-hit wonders"—seen only once.
  • The data was sparse, meaning many species were likely missed entirely.

Using MOSAIC, the authors estimated that the true number of fungal species in their study area was actually around 16,574. This suggests that about 1,200 species were present but never detected by the samplers.

They also looked at what drives these fungal communities. By feeding weather data into their model, they found that:

  • Latitude and Temperature were the biggest drivers. Fungi followed a predictable pattern: they were most diverse near the equator and less diverse as you moved toward the poles.
  • Temperature had a "hump-shaped" relationship. Fungi loved moderate warmth but struggled when it got too hot or too cold.
  • Rain had a negative effect. It seems that heavy rain washes fungal spores out of the air and onto the ground, making them harder to find in the air samples.

What This Means

The paper doesn't claim to have solved the mystery of all biodiversity. Instead, it offers a better magnifying glass. It suggests that by admitting our detection methods are imperfect and by allowing for the fact that different places have different rules, we can get a much truer picture of how life is distributed on Earth.

The authors show that their method provides a "coherent probabilistic treatment," meaning it gives us not just a single number, but a range of likely answers with a measure of confidence. When they applied it to the fungal data, the results made biological sense, confirming known patterns about how fungi react to climate.

In short, MOSAIC is a tool that helps ecologists stop guessing and start calculating. It acknowledges that the world is full of hidden guests, and it gives us a mathematical way to count them, even when they are hiding in the shadows. This is a vital step toward understanding how biodiversity changes across the planet, which is essential for protecting the natural world in a changing climate.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →