← Latest papers
💬 NLP

INS-ActBench: A Comprehensive Benchmark for Assessing Professional Actuarial Capability of Large Language Models

The paper introduces INS-ActBench, a comprehensive benchmark comprising over 12,000 questions from 16 actuarial associations to evaluate large language models across knowledge, long-context case reasoning, and tool-based practice, revealing that while frontier models excel in standardized knowledge, they still lag significantly in complex professional workflows.

Original authors: Changyu Chen, Chenwei Lin, Xian Xu

Published 2026-07-28
📖 4 min read☕ Coffee break read

Original authors: Changyu Chen, Chenwei Lin, Xian Xu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to be a financial wizard. You've already shown it how to read a dictionary of money terms and how to do basic math. But being a true wizard isn't just about knowing the spells; it's about knowing when to use them, how to mix them together when the situation is messy, and having the steady hands to actually cast the spell without setting the castle on fire. This is the world of Large Language Models (LLMs)—super-smart computer programs that can read, write, and reason like humans. For a long time, we've tested these robots on isolated tasks: "Can you define this word?" or "Can you solve this math problem?" But in the real world, professionals don't work in isolation. They have to read a hundred-page contract, find the one tiny clause that matters, calculate the risk, and then use a spreadsheet or code to prove their answer is right. This paper asks a big question: Can these AI wizards actually do the whole job, or are they just good at memorizing flashcards?

Enter INS-ActBench, a new, massive "final exam" created by researchers at Fudan University to test AI on the specific, high-stakes job of Actuarial Science. Think of an actuary as a professional risk detective who works for insurance companies. They use math, statistics, and economics to figure out how likely it is that something bad will happen (like a car crash or a house fire) and how much money the company needs to save up to pay for it. It's a job that requires knowing the rules, reading complex stories about real-world disasters, and doing precise calculations that can't be wrong. The researchers built a test bank of 12,050 questions pulled from real exams taken by actuaries around the world. They didn't just ask simple trivia; they designed the test to see if the AI could handle three different levels of difficulty:

  1. The Flashcards (INS-Act-Know): Can the AI remember the definitions and formulas?
  2. The Mystery Novel (INS-Act-Case): Can the AI read a long, complicated story about a company's finances, find the hidden clues, and make a smart guess about the future?
  3. The Workshop (INS-Act-Practice): Can the AI actually do the work by writing the correct computer code or spreadsheet formulas to get a number that can be checked?

When the researchers put nine different AI models through this grueling test, the results were a mix of "wow" and "whoops." The AI models were surprisingly good at the first level. They aced the flashcards, often scoring higher than human experts on standardized knowledge questions. It turns out that if you just ask an AI to recite a rule or solve a simple math problem, it's a genius.

However, the moment the test got messy, the robots stumbled. When the AI had to read a long, complex insurance case (the "Mystery Novel" level), their scores dropped dramatically. They struggled to connect the dots between different parts of the story. Even worse, when they had to enter the "Workshop" and actually write the code or spreadsheet formulas to produce a final number, they failed to do so reliably. The study found that while the AI could talk about the math, it often couldn't do the math in a way that was consistent and verifiable.

The researchers also discovered that the AI's performance changed depending on where the rules came from. Insurance laws are different in the US, the UK, and Japan. The AI models were great at the "standard" rules they saw often in their training, but they got confused when the questions came from different countries with different local regulations. They also found that the AI's biggest failures weren't usually because they forgot how to use a tool (like a calculator); it was because they used the tool correctly but applied the wrong logic or formula to the problem.

In short, this paper suggests that while AI has become a fantastic encyclopedia for actuarial knowledge, it is not yet a reliable professional assistant. It can pass the written test, but it isn't ready to sit in the office and make the final decisions that keep insurance companies safe. The researchers conclude that for AI to truly help in this field, it needs to get much better at connecting the dots in long stories and executing complex, step-by-step workflows without making small but costly mistakes. For now, the human expert is still the one holding the pen.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →