Model Stealing Through the Lens of Model Multiplicity
This paper challenges the assumption that high-fidelity model stealing attacks yield economically equivalent surrogates by demonstrating that the Rashomon Set of near-optimal extracted models exhibits significant diversity in critical deployment properties like fairness and robustness, despite achieving comparable accuracy to the target.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you hire a master chef to create a secret, award-winning recipe. A rival chef wants to steal that recipe, but they aren't allowed to see the ingredients or the cooking steps. Instead, they are only allowed to order dishes from the master chef, taste them, and write down the results.
Based on those taste tests, the rival chef tries to recreate the dish.
The Old Way of Thinking:
For a long time, experts believed that if the rival chef's dish tasted 99% like the master chef's dish, they had successfully stolen the recipe. They assumed the rival chef had a perfect clone of the original. If the flavors matched, the "intellectual property" was gone.
This Paper's New Perspective:
This paper argues that "tasting the same" doesn't mean "cooking the same way."
The authors suggest that there isn't just one way to recreate a dish that tastes 99% like the original. There is a whole family of recipes (which they call a "Rashomon Set") that all produce a dish tasting almost identical to the master's, but they use completely different ingredients, cooking times, or techniques.
Here is how they explain it using simple analogies:
1. The "Many Roads to Rome" Problem
Imagine you are trying to guess a secret number between 1 and 100.
- The Master Chef (Target Model): Knows the number is 42.
- The Thief (Adversary): Asks, "Is it higher than 40?" The Master says "Yes." "Is it lower than 50?" The Master says "Yes."
- The Thief's Guess: Based on this, the thief guesses 42. They are right! They stole the answer.
But what if the thief guessed 45? Or 48?
If the Master's rule was actually "Pick any number between 41 and 49," then 42, 45, and 48 are all equally correct answers. They all satisfy the "taste test" (the queries).
The paper shows that in machine learning, when a thief steals a model, they often end up with one of these "45" or "48" guesses. It tastes the same as the original (high fidelity), but if you ask a slightly different question later, or if you look at specific people, the thief's model might give a totally different answer than the original.
2. The "Shadow Puppet" Analogy
Think of the original model as a complex hand gesture making a shadow puppet of a bird on the wall.
The thief only sees the shadow (the output). They try to build a hand that makes the same bird shadow.
- The Thief's Hand: They might use a different finger arrangement, a different wrist angle, or a different hand size.
- The Result: The shadow on the wall looks exactly like a bird.
- The Catch: If the light moves slightly, or if you ask the hand to make a different shadow (one the thief never saw), the thief's hand might make a rabbit or a dog, while the original hand would still make a bird.
The paper found that even when the "shadow" (the prediction) looks perfect, the "hand" (the internal logic) is often very different.
3. The "Group Fairness" Twist
The paper also looked at how these "different hands" treat different groups of people.
Imagine the Master Chef's recipe is fair: it gives a generous portion to everyone.
The Thief's recipe tastes the same on average. But because they used a different method, they might accidentally give huge portions to one group of people and tiny portions to another group, even though the total amount of food is the same.
The study found that while the stolen models looked like perfect copies on paper, they often treated different groups of people (like different races or genders) very differently than the original model did. They redistributed the "errors" in harmful ways that the original model didn't do.
What Did They Actually Do?
The researchers didn't just steal one model. They stole a model, and then they used a trick (called "dropout," which is like randomly shaking the model's brain) to generate thousands of slightly different versions of that stolen model.
They found that:
- All of them tasted almost exactly like the original (high fidelity).
- But, when they looked closely, many of them disagreed with each other on specific details.
- Some of them were "unfair" in ways the original wasn't.
The Bottom Line
The paper concludes that stealing a model isn't as simple as making a perfect copy.
Just because a stolen model gets the right answer 95% of the time doesn't mean it is the original model. It might be a "look-alike" that works well in the lab but fails or acts unfairly in the real world.
So, if a company says, "We stole their model because our accuracy is the same," the paper says: "Not so fast. You might have stolen a different recipe entirely, and that could be dangerous."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.