← Latest papers
📊 statistics

Likelihood Inference for Latent Network Models under Snowball Sampling

This paper addresses the systematic bias in network parameter estimates caused by ignoring snowball sampling mechanisms by deriving an exact likelihood for continuous latent space models and developing a stochastic Expectation-Maximization algorithm that significantly improves inference accuracy and model fit compared to naive approaches.

Original authors: Nurzhan Sapargali, Sergio Buttazzo, Göran Kauermann

Published 2026-06-23
📖 5 min read🧠 Deep dive

Original authors: Nurzhan Sapargali, Sergio Buttazzo, Göran Kauermann

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: The "Snowball" Effect

Imagine you want to understand the entire social circle of a massive city, but you can't talk to everyone. Instead, you start with one person (the "seed") and ask, "Who are your friends?" Then you ask those friends, "Who are your friends?" and so on.

This is called Snowball Sampling. It's like rolling a snowball down a hill; it starts small but picks up more and more snow (people) as it goes.

The Trap:
The problem is that this method is biased. People with lots of friends (popular hubs) are much more likely to get picked up by the snowball than people with only one or two friends. If you just look at the people you collected and try to guess how the whole city is connected, you will get it wrong. You will think the city is much more tightly connected and "clumpy" than it actually is, because you've accidentally gathered mostly the popular crowd.

The Solution: A New Mathematical Lens

The authors of this paper developed a new way to do the math so that researchers can fix this bias. They created a "corrected lens" that accounts for the fact that the snowball sampling method naturally misses the quiet, isolated people.

Here is how they did it, broken down into three parts:

1. The "Invisible Wall" Analogy

Imagine the people you didn't find are standing behind an invisible wall.

  • The Naive Mistake: If you ignore the sampling method, you assume everyone you didn't find is just like the people you did find. You assume they are all hanging out in the middle of the party.
  • The Correction: The authors realized that if a person wasn't found by the snowball, it means they didn't have a connection to the people you did find in the earlier rounds. They are "invisible" because they are too far away from the main group.
  • The Math Magic: The paper proves that for a specific class of models (called Continuous Latent Space models), you can calculate the probability of this "invisibility" without needing to see the missing people. It's like knowing that if you didn't hear a noise, the person making it must be very far away, even if you can't see them. This allows them to write a perfect mathematical formula (a likelihood) that includes the missing people without actually having to interview them.

2. The "Hidden Map" (Latent Space)

To understand the network, the authors imagine every person has a secret coordinate on a hidden map.

  • If two people are close together on this hidden map, they are likely to be friends.
  • If they are far apart, they probably aren't.

The Naive Approach: When researchers ignore the sampling bias, they draw this hidden map too small. They squish everyone together into a tiny, crowded room. Because everyone is so close on this tiny map, the math predicts that everyone should be friends with everyone else. This leads to a prediction that the network has way more connections than it actually does.

The Corrected Approach: The new method stretches the map out. It realizes, "Ah, the people we missed are actually very far away on the map, which is why the snowball didn't roll over to them." This creates a much larger, more accurate map where the distances between people reflect reality.

3. The Real-World Test: The Patent Inventors

To prove their method works, the authors tested it on a real, massive network: German semiconductor inventors (about 6,000 people who work together on patents).

  • The Experiment: They took 500 different "snowballs" (samples) from this network and tried to reconstruct the whole picture using two methods: the old "naive" way and their new "corrected" way.
  • The Results:
    • The Naive Way: It drew a map where everyone was squished together. It predicted the network had nearly twice as many connections as it actually did. It thought the world was much smaller and more connected than it is.
    • The Corrected Way: It drew a map that was 8 times wider (more spread out). This matched the real data much better. It correctly predicted the number of connections and the "shape" of the network.

Why This Matters

The paper shows that if you use the old, naive method, you might draw the wrong conclusions.

  • Example: If you are studying how geography affects collaboration, the naive method might say, "Oh, distance doesn't matter much because everyone is so close on our tiny map."
  • The Truth: The corrected method shows that distance does matter a lot. The further apart two inventors are, the less likely they are to work together. The naive method hides this reality by squishing the map.

The Bottom Line

The authors didn't just say, "Snowball sampling is biased." They built a specific mathematical tool that fixes that bias for a wide range of network models. They showed that by accounting for the "invisibility" of the people you missed, you can get a true picture of the network, even when you can't talk to everyone.

In short: Don't just look at the snowball you rolled; use their new math to figure out who was left behind in the snow, so you can draw the map of the whole world correctly.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →