← Latest papers
💬 NLP

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks

This paper introduces ActuBench, a multi-agent LLM pipeline that automates the generation and evaluation of advanced actuarial reasoning tasks aligned with the IAA syllabus, demonstrating that independent verification significantly improves item quality while revealing that locally-hosted open-weights models offer superior cost-performance trade-offs and that open-ended evaluation is essential for accurately discriminating top-tier model capabilities.

Original authors: Jan-Philipp Schmidt

Published 2026-04-23
📖 5 min read🧠 Deep dive

Original authors: Jan-Philipp Schmidt

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to build a massive, ultra-difficult exam for the world's most precise accountants and risk managers: Actuaries. These are the math wizards who calculate the odds of everything from car crashes to life expectancy to set insurance prices.

The problem is, writing these exams is incredibly hard, expensive, and slow. You need human experts to write questions that are tricky but fair.

Enter ActuBench. Think of this not as a single person writing a test, but as a high-tech, automated factory run by a team of AI robots working together to build, check, and grade these exams.

Here is how the paper explains this system, broken down into simple concepts:

1. The Factory Floor: A Team of Specialized Robots

Instead of asking one AI to do everything (which often leads to mistakes), the authors built a pipeline with four distinct "agents" (AI roles), each with a specific job:

  • Agent A (The Architect): This is the smartest, most expensive robot. Its job is to look at a learning goal (like "calculate pension risk") and draft the actual exam question. It's the creative writer.
  • Agent B (The Trickster): Once the question is written, Agent B's job is to invent the "wrong answers" (distractors). These need to look plausible enough to fool a student who is halfway there, but clearly wrong to an expert.
  • Agent C (The Strict Inspector): This is the most important innovation. Agent C never writes anything. It only looks at what A and B produced. It acts like a strict teacher grading a draft. If the question is confusing or the wrong answers are too obvious, Agent C sends it back for a fix.
    • The Analogy: Imagine a writer (A) and a joke-writer (B) trying to make a comedy sketch. Agent C is the director who says, "That joke doesn't land," or "The setup is unclear." Because C didn't write the sketch, it doesn't have "blind spots" and can spot errors the writers missed.
  • The Assistant (The Librarian): A cheap, fast robot that fetches information from Wikipedia to make sure the facts are real and tags the questions by topic (e.g., "Health Insurance" or "Pensions").

The Magic: This "Inspector" role is the paper's big breakthrough. By separating the writer from the checker, they found that the AI caught 60% of the mistakes on the very first try, fixing them automatically without a human needing to step in.

2. The Two Types of Exams

The researchers didn't just test the AI on one kind of question. They ran two different types of tests to see how the AI really thinks:

  • The Multiple-Choice Test (MCQ): The AI sees a question with four options (A, B, C, D) and picks one.
    • The Catch: This is like a "recognition" test. Even if the AI doesn't know the answer, it might guess correctly by eliminating the silly options.
  • The "Judge" Test (Open-Ended): The AI has to write out the answer and explain its reasoning from scratch, with no options to choose from.
    • The Catch: This is a "derivation" test. The AI can't guess; it has to actually do the math and logic.

3. The Big Surprises (The Results)

The team tested 50 different AI models (from big companies like Google, OpenAI, and Anthropic, plus some open-source ones). Here is what they found:

  • Surprise #1: The "Cheap" Winners.
    You might think the most expensive, super-smart AI models would win easily. But on the multiple-choice test, a few models tied for first place (98% accuracy). However, the cost to run them varied wildly.

    • The Analogy: It's like two cars finishing a race in the same time. One is a $200,000 Ferrari (expensive AI), and the other is a $20,000 Toyota (open-source AI). The Toyota got you to the finish line for a fraction of the price.
    • The Winner: A model running on a standard home computer (Gemma) or a cheap cloud service (Cerebras) performed almost as well as the most expensive ones.
  • Surprise #2: The "Thinking" Mode isn't Always Worth It.
    Some companies offer a "Reasoning Mode" where the AI takes extra time to "think" before answering.

    • The Result: On multiple-choice questions, this extra thinking only gave a tiny boost (a few percentage points) but cost 2x to 5x more money. It was like paying for a luxury upgrade that didn't actually make the car go much faster.
  • Surprise #3: The "Multiple Choice" Trap.
    This is the most important finding. The AI models looked very similar on the Multiple-Choice test (all scoring near 100%). But when they took the Open-Ended "Judge" test, the rankings changed completely.

    • The Analogy: Imagine a student who memorizes the answer key. On a multiple-choice test, they get an A+. But on a test where they have to show their work, they fail.
    • The "Multiple Choice" format inflated the scores, hiding who was actually good at deep reasoning. Only the "Judge" test revealed the true difference between the top-tier AI and the rest.

4. Why This Matters

The authors built a public website where anyone can look at these questions and see how different AIs answered them.

The Takeaway:
If you just need an AI to pick the right answer from a list (like a quick check), you don't need the most expensive, "thinking" AI. You can use a cheaper, open-source model running on your own computer and get great results.

But, if you need the AI to solve a complex problem from scratch (like a real-world actuarial calculation), you do need the top-tier, expensive "reasoning" models. The cheap models can recognize the answer, but they struggle to derive it without help.

In short: ActuBench is a new, automated factory that builds better exams for AI, proving that while cheap AI is great at guessing, the expensive "thinking" AI is still the only one that can truly solve the hardest problems.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →