← Latest papers
🤖 machine learning

Who Trains Matters: Federated Learning under Enrollment and Participation Selection Biases

This paper addresses the persistent performance gap in federated learning caused by both enrollment and participation selection biases by formalizing a two-stage selection model and proposing \textsc{FedIPW}, an inverse-probability-weighted aggregation scheme that effectively recovers target-population objectives even when client-level covariates are limited.

Original authors: Gota Morishita

Published 2026-04-30
📖 6 min read🧠 Deep dive

Original authors: Gota Morishita

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to bake the perfect cake for a whole city. To do this, you ask thousands of home bakers to send you a small piece of their batter so you can mix it together and figure out the ideal recipe. This is essentially how Federated Learning (FL) works: instead of gathering all the data in one place, a central server asks many devices (like phones) to train a model locally and send back only the "updates" (the batter pieces).

The problem, as this paper explains, is that who you ask to send batter matters just as much as how you mix it.

The Two-Stage Filter: Who Gets in the Door?

The paper argues that in real life, the bakers you end up hearing from are rarely a perfect cross-section of the whole city. This happens in two distinct stages, like a two-step security check at a concert:

  1. The "Enrollment" Bias (Who gets the invitation?):
    First, you have to be eligible to join the project. Maybe you need a specific type of phone, a certain software version, or you just have to click "I Agree" on a consent form. If your phone is old or you live in an area with poor internet, you never even get the invitation. You are filtered out before the game even starts. The paper calls this Enrollment Bias.

    • Analogy: Imagine you only invite people who own a red car to your baking club. Even if you ask everyone with a red car to participate, you've already missed everyone with a blue car, a bike, or no vehicle at all. Your "baking club" is already skewed.
  2. The "Participation" Bias (Who actually shows up?):
    Second, even among the people who did get invited, not everyone shows up to every meeting. Maybe their battery is dead, their internet is spotty, or it's 3 AM in their time zone. They are enrolled, but they don't participate in that specific round. The paper calls this Participation Bias.

    • Analogy: Even if you invited everyone with a red car, maybe only the ones who are awake and have a full tank of gas actually drive to the meeting.

The Problem: Baking the Wrong Cake

Most existing methods try to fix the second problem (who shows up). They say, "Okay, the people who showed up tonight are mostly night-shift workers; let's adjust the recipe to account for that."

But this paper points out a bigger issue: If the people who got invited in the first place (the red car owners) aren't like the rest of the city, fixing the "who shows up" part won't help. You might perfectly adjust for the night-shift workers, but you are still baking a cake based entirely on red-car owners. The final result will taste great to red-car owners but terrible to everyone else.

The paper calls this a "Target-Population Mismatch." The model learns to serve the people who are reachable, not the people it's supposed to serve.

The Solution: A Weighted Scale (FedIPW)

To fix this, the author proposes a new method called FedIPW (Federated Inverse Probability Weighting).

Think of this as using a weighted scale instead of a simple average.

  • The Old Way (FedAvg): If 10 people send updates, you give each person 1/10th of the weight.
  • The New Way (FedIPW): You look at who didn't show up and ask, "Why?"
    • If a group of people (say, people with Android phones) rarely gets invited because of strict software rules, but they do show up when they are invited, the algorithm gives their updates extra weight.
    • If a group (say, people with new iPhones) is invited often but rarely shows up, their updates are also weighted carefully to represent the ones who did show up.

By mathematically "re-weighting" the updates based on the probability of being invited and the probability of showing up, the server can reconstruct what the "average city baker" would have contributed, even if they never actually sent a piece of batter.

What If We Don't Have All the Details? (The "Limited Information" Fix)

Sometimes, the server doesn't know the details of the people who didn't get invited (e.g., it doesn't know how many people in the city have old phones). It only knows the big picture (e.g., "20% of the city uses Android").

In this case, the paper suggests a Calibration trick.

  • Analogy: Imagine you are baking with a sample of bakers, but you don't know the exact demographics of the whole city. However, you do have a census report that says, "The city is 50% men and 50% women."
  • If your sample of bakers is 80% men, you can't just ignore the women. Instead, you give the men's updates less weight and the women's updates more weight until your sample looks like the census report (50/50).
  • This doesn't fix everything perfectly, but it gets you much closer to the right recipe than doing nothing.

The "Bias Floor" Warning

The paper also warns about a "Bias Floor."
Imagine you are trying to hit a bullseye. If you miss the target slightly because your aim is shaky (random error), you can get better with practice. But if your gun is bent (structural error), you will always miss the center, no matter how much you practice.

The paper proves that if you ignore the "Enrollment" stage (the bent gun), you will hit a Bias Floor. No matter how many rounds of training you do, the model will never reach the true best solution for the whole population. It will get stuck in a "good enough" zone that is actually wrong for the people you care about.

Summary

  • The Issue: Federated learning often fails because the people who join the training (enrollment) and the people who actually participate are not representative of the whole population.
  • The Fix: Use a two-step correction (FedIPW) that mathematically weighs the updates to account for who was excluded at the start and who dropped out during the process.
  • The Takeaway: It's not enough to just fix who shows up to the meeting; you have to fix who was invited to the meeting in the first place. If you don't, your model will be biased toward a specific group, no matter how smart the algorithm is.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →