← Latest papers
💰 quantitative finance

Discrimination-free Insurance Pricing with Privatized Sensitive Attributes

This paper proposes statistical methods for estimating discrimination-free insurance premiums using only privatized sensitive attributes held by a trusted third party, providing theoretical guarantees and empirical validation for fair pricing under both known and unknown privacy mechanisms.

Original authors: Tianhe Zhang, Suhan Liu, Peng Shi

Published 2026-06-16
📖 5 min read🧠 Deep dive

Original authors: Tianhe Zhang, Suhan Liu, Peng Shi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to set the price for a car insurance policy. To do this fairly and accurately, you need to know how much risk a driver brings. Usually, you look at things like their age, driving history, and the type of car they drive.

But there's a problem: In many places, the law says you cannot use certain personal details to set prices, like a person's gender or race. This is to prevent discrimination. However, there's a catch: even if you don't ask for gender, your computer model might still figure it out by looking at other clues (like zip codes or shopping habits) that are strongly linked to gender. This is called "indirect discrimination."

Recent experts have come up with a mathematical recipe called a "Discrimination-Free Premium." This recipe tries to calculate a price that ignores gender completely, ensuring that two people with the same driving habits pay the same, regardless of their gender.

The Big Problem:
To use this "Discrimination-Free" recipe, you usually need to see the actual gender data to do the math. But, because of privacy laws, insurance companies often aren't allowed to see the real gender data at all. It's like trying to bake a cake using a recipe that requires eggs, but the store only sells you "egg-flavored powder" that looks like eggs but isn't quite the same.

The Solution in This Paper:
The authors of this paper (Tianhe Zhang, Suhan Liu, and Peng Shi) figured out how to bake that cake using only the "powder." They created a new method that works even when the sensitive information (like gender) has been "scrambled" or "noised up" to protect privacy.

Here is how they did it, using some simple analogies:

1. The "Trusted Neighbor" Setup

Imagine a three-person team:

  • The Insurance Company: They have the customer's driving data (age, car type) and the claim history (how much money was paid out). They don't have the gender data.
  • The Trusted Third Party (TTP): A neutral, secure vault that holds the scrambled gender data.
  • The Customer: They provide their data.

The Insurance Company sends their driving data to the Trusted Neighbor. The Neighbor mixes it with the scrambled gender data to do the math, then sends back the final "fair price" without ever revealing who is actually male or female.

2. The "Scrambled Egg" (Privatized Data)

The paper deals with two scenarios regarding how "scrambled" the data is:

  • Scenario A: We know the scramble recipe.
    Imagine the Trusted Neighbor tells the Insurance Company, "I added exactly 10% noise to the gender data." The authors show that if you know exactly how the data was scrambled, you can mathematically "unscramble" the bias and calculate the fair price perfectly. It's like knowing exactly how much salt was added to a soup, so you can calculate how much water to add to fix the taste.

  • Scenario B: We don't know the scramble recipe.
    This is harder. Imagine the Neighbor says, "I scrambled it, but I forgot exactly how much noise I added." The authors developed a clever trick to guess the noise level using a special "anchor point" in the data (a specific type of driver where we are 100% sure of their gender). Once they estimate the noise level, they can still calculate a fair price, though it's slightly less precise than if they knew the exact recipe.

3. The "Magic Filter" (Transformed Data)

The paper also introduces a cool trick. Before sending data to the Trusted Neighbor, the Insurance Company runs the driving data through a "magic filter" (a neural network). This filter learns to highlight the parts of the data that actually predict risk (like "smoking status" or "age") while smoothing out the noise.

  • The Result: Even if the gender data is very noisy (very scrambled), using this "magic filter" on the driving data helps the model stay accurate. It's like wearing noise-canceling headphones; even if the room is loud, you can still hear the music clearly.

What the Experiments Showed

The authors ran computer simulations and tested their method on real-world health and car insurance data. They found:

  • Fairness Works: Their method successfully removed both direct and indirect discrimination.
  • Privacy is Safe: The insurance company never needed to see the real gender data.
  • Noise Matters: If the privacy "noise" is too high (too much scrambling), the prices get less accurate. However, their method handles this better than older methods.
  • Underestimating is Risky: If you guess the noise level is lower than it actually is, your prices get messed up more than if you guess it's higher. It's better to be slightly "conservative" in your guess.

The Bottom Line

This paper provides a practical toolkit for insurance companies. It allows them to use advanced AI to set fair prices that don't discriminate based on gender or race, even when strict privacy laws prevent them from ever seeing the real gender data. They can do this by working with a trusted third party and using statistical tricks to correct for the "noise" added to protect privacy.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →