← Latest papers
📊 statistics

Statistics-Friendly Confidentiality Protection for Establishment Data, with Applications to the QCEW

This paper proposes a novel, policy-friendly confidentiality framework for business data that shifts focus from membership inference to attribute inference protection by defining uncertainty intervals around establishment values, and validates this approach using both confidential U.S. Bureau of Labor Statistics data and a publicly released substitute dataset.

Original authors: Kaitlyn Webb, Prottay Protivash, John Durrell, Daniell Toth, Aleksandra Slavković, Daniel Kifer

Published 2026-03-23
📖 6 min read🧠 Deep dive

Original authors: Kaitlyn Webb, Prottay Protivash, John Durrell, Daniell Toth, Aleksandra Slavković, Daniel Kifer

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a statistician working for the government. Your job is to publish a giant report about how many people are working and how much money they are making across the country. This report is called the QCEW (Quarterly Census of Employment and Wages).

The problem? You can't just publish the raw numbers. If you say "The Walt Disney Company in California has 36,000 employees," that's fine. But if you say "Bob's Ice Cream Shop in a small town has 3 employees," and you publish that exact number, you might accidentally reveal private information about Bob's employees.

For decades, statisticians have used a method called "Cell Suppression" to protect this data. Think of this like a game of "Redact the Secret." If a number is too revealing, they simply black it out (suppress it) so no one can see it.

  • The Problem: This is like trying to solve a puzzle but throwing away 60% of the pieces. The resulting picture is blurry and useless for economists who need to see the big trends.

The New Idea: "The Foggy Window"

This paper proposes a smarter way to protect business data. Instead of blacking out numbers, they want to add a little bit of "statistical fog" (noise) to the answers. This way, the numbers are still visible, but they are slightly fuzzy, making it impossible to know the exact truth while still knowing the general truth.

However, there's a catch: Businesses are not all the same size.

  • Small Businesses: A tiny shop with 3 employees. If you add a little fog, it might look like they have 2 or 4 employees. That's a huge change! We need strong fog here to protect them.
  • Big Businesses: A giant company with 36,000 employees. If you add the same amount of fog, it might look like they have 36,002 employees. That's a tiny change. But if you make the fog too thick to protect them, the whole national economy looks wrong.

The Old Way vs. The New Way

The Old Way (Personalized Privacy):
Previous attempts tried to give every business its own "privacy setting."

  • The Logic: "Big companies are less private, so we give them weak protection. Small companies are very private, so we give them strong protection."
  • The Problem: This feels unfair to policymakers. It's like saying, "We will protect your neighbor's house with a steel door, but your house gets a screen door because you are richer." It creates a two-tier system of privacy that is hard to justify.

The New Way (Gaussian Establishment DP):
The authors propose a new framework called Gaussian Establishment DP (let's call it G-EDP). Instead of changing the strength of the protection based on size, they change the type of protection.

The Analogy: The "Fuzzy Target" vs. The "Ruler"

Imagine you are trying to guess the weight of a person.

  1. For a small business (The "Ruler" approach): You need to be very careful about the percentage error. If a shop has 10 employees, guessing 11 is a 10% error. That's bad. We need to make sure the "fog" covers a wide range relative to their size.
  2. For a big business (The "Fuzzy Target" approach): You don't care about the percentage as much; you care about the absolute number. If a giant company has 100,000 employees, guessing 100,100 is only a 0.1% error. That's fine! But guessing 100,000 exactly is bad.

The paper introduces a mathematical tool called a "Neighbor Function" (specifically, the Square Root function).

  • Think of this as a special lens. When you look at a small number through this lens, it gets magnified (making the "fog" cover a larger relative area). When you look at a huge number, the lens compresses it (making the "fog" cover a smaller relative area, but a larger absolute area).
  • The Result: Small businesses get a wide "uncertainty zone" (e.g., "We know they have between 2 and 5 employees"). Big businesses get a wide "uncertainty zone" too, but in absolute terms (e.g., "We know they have between 35,900 and 36,100 employees").
  • The Win: Everyone gets a fair "zone of plausible deniability." No one is strictly "less protected" than the other; they just have different kinds of protection that fit their size.

The Two Magic Machines

The paper builds two "machines" (algorithms) to generate these fuzzy answers:

  1. The ψ\psi-Machine (The "Black Box"):
    This machine adds the fog perfectly according to the rules. However, it has a quirk: it doesn't tell you exactly how much fog it added. It's like a chef who adds a secret spice but won't tell you the recipe. It's great for privacy, but statisticians hate it because they can't calculate the "margin of error" for their reports.

  2. The PNC-Machine (The "Honest Chef"):
    This is the paper's big innovation. It uses a clever trick called "Probably No Clipping."

    • The Trick: Before adding the fog, the machine looks at the data and says, "I bet no single business is bigger than XX." It sets a "cap" based on this guess.
    • Because it knows the cap, it can calculate exactly how much fog to add and tell you the exact margin of error.
    • Why it matters: Statisticians can now say, "This number is 10,000, give or take 500, and we are 99% sure that's accurate." This makes the data actually useful for making economic decisions.

Why This Matters

  • For the Government: They can release more data without fear of exposing secrets. No more blacking out 60% of the report!
  • For Economists: They get clearer pictures of the economy. They can see trends in small towns and big cities without the data being distorted by huge, random errors.
  • For Policy Makers: They get a system that treats all businesses fairly. A small bakery and a giant tech firm both get "plausible deniability"—a safe zone where the exact truth is hidden, but the general truth is clear.

In Summary:
The paper solves the problem of "How do we hide the exact number of employees in a business without making the whole economic report useless?" by inventing a new kind of "smart fog" that adjusts its thickness based on the size of the business, ensuring everyone gets a fair share of privacy while keeping the data useful for everyone else.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →