← Latest papers
💻 computer science

Validating Large Language Model Extraction of Actuarial Variables from Unstructured Claims Documents

This study evaluates a large language model pipeline for extracting 14 actuarial variables from workers' compensation claims, revealing moderate overall agreement with human reviewers (quadratic weighted kappa of 0.53) and significant variability across dimensions, which underscores the need for calibration and phased deployment before formal validation.

Original authors: Robert Lieberthal, Vietbao Phan, Jawand Singh, Elizabeth Sottung, Richard Tran

Published 2026-06-25
📖 5 min read🧠 Deep dive

Original authors: Robert Lieberthal, Vietbao Phan, Jawand Singh, Elizabeth Sottung, Richard Tran

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a massive library of insurance claim files. These files are messy: they contain handwritten doctor notes, phone call transcripts, legal letters, and medical records. For decades, insurance experts (actuaries) have had to read these messy files by hand to figure out how much money they might need to set aside for future payouts. It's slow, expensive, and prone to human error.

This paper asks a simple question: Can a super-smart computer (a Large Language Model, or LLM) read these messy files and pull out the specific numbers and facts actuaries need, just as well as a human can?

Here is the breakdown of their experiment, explained simply:

The Setup: A "Fake" Library Test

The researchers didn't want to risk testing on real people's private medical data, so they built a simulated library. They used a computer program to generate 20 "fake" insurance claims. Because the computer created them, the researchers knew the "correct" answers (the ground truth) for every single fact in those files.

They fed these fake files into their AI system. The AI was tasked with pulling out 14 specific pieces of information, like "How bad was the injury?" or "Is there a risk of a lawsuit?"

The Judges: The Human Reviewers

To see if the AI was doing a good job, they hired two expert human reviewers (pharmacists with advanced degrees) to grade the AI's work. They used a report card with 10 different categories, such as:

  • Relevance: Did the AI talk about the right thing?
  • Accuracy: Did it get the facts right?
  • Hallucination: Did the AI make up facts that weren't in the file?
  • Missing Info: Did the AI forget to mention something important?

If the two human judges disagreed by a lot (more than 2 points on a 5-point scale), a third judge stepped in to make the final call.

The Results: The "Good, The Bad, and The Weird"

1. The Overall Score: "Okay, but not perfect"
The two judges agreed with each other about 51% of the time on the exact same score. When you account for them being "close" (within one point), they agreed 81% of the time. In the world of statistics, this is considered "moderate" agreement. It's like two weather forecasters agreeing it will rain, but one says "light drizzle" and the other says "heavy storm."

2. The Good News: The AI is great at the basics
The judges agreed almost perfectly on whether the AI was relevant (talking about the right topic) and appropriate (writing clearly). If the AI said, "The patient broke a leg," the judges agreed that was the right topic.

3. The Bad News: The AI struggles with "What's Missing" and "What's Fake"
This is where the judges fought the most.

  • Missing Information: One judge might think the AI missed a crucial detail, while the other thinks it got everything. They disagreed 40% of the time. It's like two people looking at a puzzle; one thinks a piece is missing, and the other thinks the picture is complete.
  • Hallucinations: The AI sometimes made things up. The judges had a hard time agreeing on how bad the made-up facts were.
  • Context: The AI sometimes misunderstood medical jargon, and the judges couldn't agree on how confused the AI actually was.

4. The "Grumpy" Judge
One of the human judges was consistently stricter than the other. When they disagreed, the third judge (the tie-breaker) sided with the "looser" judge (the one giving higher scores) 10 out of 15 times. This suggests that the "stricter" judge was being overly cautious, perhaps too worried about missing things that weren't actually missing.

What Does This Mean for the Future?

The researchers draw a clear line in the sand based on their findings:

  • The "Triage" Zone (Safe for now): The AI is reliable enough to be used as a sorter. Imagine a mailroom where the AI scans a pile of letters and says, "Hey, this one looks urgent, put it on the top of the human's desk." The AI is good at spotting the obvious stuff.
  • The "Money" Zone (Not ready yet): The AI is not ready to be used for the final calculation of how much money the insurance company needs to save. Because the AI sometimes misses details or invents facts, and because humans can't even agree on whether those mistakes happened, you can't trust the AI to do the final math on its own yet.

The Bottom Line

The paper concludes that we need to calibrate our human judges better before we trust the AI. Right now, the AI is a helpful intern who can do the heavy lifting of reading the files, but a senior expert still needs to double-check the final numbers, especially to make sure the intern didn't miss anything or make anything up.

The researchers also noted that the AI's "confidence score" (a number it gives itself saying "I'm 90% sure!") wasn't very useful because it was always high, even when it was wrong. It's like a student who always says "I'm sure I got an A" even when they got a C.

In short: The technology is promising and works well for sorting and flagging claims, but it needs more training and better human oversight before it can be trusted to handle the actual money calculations.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →