On the Fragility of Data Attribution When Learning Is Distributed
This paper demonstrates that data attribution in distributed learning is fragile, as a malicious participant can exploit latent optimization to inject synthetic batches that significantly inflate their measured contribution without degrading global model utility or triggering existing defenses.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: The "Credit Card" Problem
Imagine a group of neighbors trying to build a giant, shared garden (the Machine Learning Model). Each neighbor brings a different set of seeds and tools (their Data). Some neighbors have rare, exotic flowers; others only have common weeds.
To keep everyone motivated, the group leader uses a special calculator called Data Attribution. This tool tries to figure out exactly how much each neighbor contributed to the garden's final beauty. Based on this score, neighbors get paid, get credit, or get to keep their spot in the club.
The paper's main discovery: A sneaky neighbor can trick this calculator. They can make it look like they brought the most valuable seeds in the whole group, even though they didn't actually help the garden grow any better. In fact, the garden looks exactly the same as if they had played fair.
The Setup: How the Attack Works
The researchers found a way for a single "bad actor" to game the system without getting caught. Here is how they did it, broken down into steps:
1. The "Ghost Seeds" (Synthetic Data)
Usually, if you want to cheat, you might bring fake, broken seeds (bad data) to ruin the garden. But that's obvious; the garden would look ugly, and you'd get kicked out.
Instead, this attacker uses Latent Optimization. Think of this as a "magic seed printer." The attacker has a blueprint (a decoder) that can print tiny, perfect-looking seeds. These aren't real seeds from their own land; they are generated by a computer.
2. The "Missing Puzzle Piece" Strategy
The garden is missing some specific types of flowers because the other neighbors didn't bring them. The attacker's magic printer creates just enough of these missing flowers to fill the gaps.
- Why this matters: The credit calculator (the attribution tool) loves "completeness." It thinks, "Wow, this neighbor filled the holes in our garden! They must be super helpful!"
- The trick: The attacker only prints just enough to look helpful, but not enough to mess up the garden's overall look.
3. The "Perfect Mimic" (Stealth)
To avoid getting caught, the attacker makes sure their contribution looks exactly like a normal, honest neighbor's contribution.
- They match the size of the contribution (so it doesn't look too big).
- They match the direction (so it pushes the garden in the same way everyone else does).
- They make sure the final garden (the model's accuracy) looks just as beautiful as it would have without them.
The Result: The "Invisible" Heist
The paper ran this experiment with different types of gardens (datasets like CIFAR-10 and FashionMNIST) and different garden planners (models like ResNet and VGG).
What happened?
- The Score: The sneaky neighbor's "contribution score" skyrocketed. They went from being at the bottom of the list to the top, or at least near the top.
- The Garden: The garden's quality (accuracy) did not drop. It stayed exactly the same.
- The Defenses: The group's security guards (defenses that look for weird shapes or broken plants) didn't notice anything wrong because the "ghost seeds" looked so normal.
The Analogy: The "Perfectly Polished" Lie
Imagine a team of chefs making a soup. The boss asks, "Who added the most flavor?"
- Normal Chefs: Add real ingredients.
- The Attacker: Instead of adding a huge, obvious pile of salt (which would ruin the soup), they add a tiny, invisible pinch of a "flavor enhancer" that the boss's taste-tester loves.
- The Outcome: The soup tastes exactly the same as before (no one complains), but the taste-tester's machine gives the attacker a massive bonus because the machine thinks that tiny pinch was the secret ingredient that made the soup perfect.
Why This Matters (According to the Paper)
The paper warns us that trust is fragile.
- We are starting to use these "contribution scores" to decide who gets paid for data, who owns the model, and how to govern AI systems.
- The paper shows that these scores can be manipulated easily. A bad actor can steal credit without hurting the system's performance.
- Current security measures (which check if the model is broken or if the data looks weird) do not work against this specific type of trick.
Summary
The paper proves that in a distributed learning system, you cannot trust the "scorecard" just because the final result looks good. A clever participant can use a "magic printer" to create fake-but-perfect data that tricks the scoring system into giving them a massive reward, all while leaving the actual product unchanged and undetectable.
The takeaway: If you are paying people based on how much they "contributed" to an AI, you need a new way to check their work, because the current scorecards can be fooled.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.