← Latest papers
📊 statistics

Severity estimation in dependent collective risk models

This paper demonstrates that fitting severity models directly to pooled claims in dependent collective risk settings yields inconsistent estimates due to size-bias, and proposes a consistent composite likelihood estimation procedure based on the correct observed-claim distribution, establishing its asymptotic properties and validating its performance through simulations.

Original authors: Christopher Blier-Wong

Published 2026-07-10
📖 5 min read🧠 Deep dive

Original authors: Christopher Blier-Wong

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are running a massive lemonade stand. Every day, you track two things: how many cups you sell (the count) and how much sugar is in each cup (the severity). In the old, simple way of thinking, you assume these two things are totally unrelated. You think, "Okay, I sold 5 cups today, and I sold 100 cups yesterday; the sugar in those cups is just a random mix of all the sugar I ever used."

But what if they aren't random? What if, on the days you sell a ton of lemonade, you accidentally make the cups extra sweet? Or maybe on slow days, you only make small, weak cups? This is the real-world mess the paper tackles: frequency and severity are often linked.

The Great Sampling Mix-Up

Here is the big problem the authors found: When you look at all the lemonade cups you sold over a year to figure out your "average sugar level," you are actually looking at a biased sample.

Think of it like this:

  • Customer A buys 1 cup.
  • Customer B buys 100 cups.
  • Customer C buys 0 cups (they just looked at the menu and left).

If you dump every single cup from every customer into one giant bucket to measure the sugar, Customer B is dominating the bucket. They contributed 100 times more data than Customer A. If Customer B happened to buy the super-sweet cups, your giant bucket will taste way too sweet, even if the average cup you could have sold was much less sweet.

The paper proves mathematically that if you just grab all the cups from the bucket and try to guess the "true" sugar recipe, you will get it wrong. You are measuring the size-biased mix (the bucket's taste), not the marginal recipe (the true sugar level of a single, random cup).

What the Paper Says You Shouldn't Do

The authors explicitly warn against the "naive" approach. This is when an insurance company (or lemonade stand) says, "Let's just take all the claims we received, ignore who bought them, and fit a curve to them."

They show that this method is inconsistent. It doesn't matter if you have a million customers or a billion; if you ignore the link between how many claims a person has and how big those claims are, your estimate of the "true" severity will always be off. It's like trying to guess the average height of all people in a city by only measuring the people standing on the tallest building's observation deck. You'll get a number, but it won't be the truth.

The New, Smarter Way: The Composite Likelihood

So, how do we fix the bucket? The authors built a new tool called a Composite Likelihood.

Instead of ignoring the bias, they embrace it. They realized that the bucket does contain a specific, predictable pattern: it's a "size-biased mixture." It's not random noise; it's a specific mathematical distortion.

Their new method works in two clever steps:

  1. Count the Customers: First, they look at how many people bought lemonade (the frequency).
  2. Taste the Bucket Correctly: Then, they look at the giant bucket of cups, but they use a special formula that says, "Okay, we know Customer B is over-represented. Let's mathematically adjust for that to figure out what the true sugar recipe must have been."

They call this a "composite" method because it stitches together two different pieces of the puzzle (the count and the adjusted bucket taste) rather than trying to solve the whole impossible jigsaw at once.

How Sure Are They?

The authors didn't just guess this would work; they ran simulations to prove it.

  • They created a fake lemonade world with 1,000,000 customers.
  • They made the "naive" method try to guess the sugar level. It was off by about 25% (a huge error!).
  • Then they used their new "composite" method. It nailed the true sugar level, with almost zero bias.

They also tested how confident they should be in their numbers. In statistics, when you have groups of data (like 100 cups from one customer), you can't treat every cup as a totally independent fact. The authors showed that if you treat them as independent, you get false confidence (your error bars look too small). Their new method uses something called "Godambe information" (a fancy sandwich of math) to make sure the error bars are wide enough to be honest.

The "Sarmanov" Secret Sauce

To make all this math work in the real world, they used a specific type of mathematical model called the Sarmanov copula. Think of this as a special kind of glue that lets them stick the "number of cups" and the "sugar level" together in a way that is flexible but still keeps the math solvable.

Because this glue has a special structure, they could write down the exact formula for the "biased bucket" taste. This allowed them to build their correction tool without needing to know every single tiny detail of the universe.

The Bottom Line

If you are an insurance company or a risk manager:

  • Don't just pool all your claims together and fit a curve. You will be wrong.
  • Do realize that customers with many claims are skewing your data.
  • Use the new method that corrects for this skew.

The paper shows that by understanding how the data is collected (the size-bias), we can fix the recipe. It's not magic; it's just realizing that the bucket doesn't taste like the recipe, and knowing exactly how to translate the bucket's taste back into the truth.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →