← Latest papers
🤖 AI

FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables

This paper introduces FinProBench, a benchmark for evaluating financial AI agents, and Role-Grounded Rubric Construction (RGRC), a pipeline that derives evaluation criteria from professional deliverables to effectively capture tacit standards and significantly outperform prompt-based methods in role-specialized financial tasks.

Original authors: Ben Wang, Kang Zhou, Lifan Guo, Feng Chen, Chi Zhang

Published 2026-08-06
📖 3 min read☕ Coffee break read

Original authors: Ben Wang, Kang Zhou, Lifan Guo, Feng Chen, Chi Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to be a professional, like a doctor, a lawyer, or a financial analyst. In the world of artificial intelligence, we have these super-smart "agents" that can chat, write code, and solve puzzles. But how do we know if they are actually doing a good job, or just sounding like they are? This is the big question in the field of AI evaluation. For a long time, we've tested AI by asking it simple questions with right-or-wrong answers, like a multiple-choice quiz. But real professional work isn't a quiz; it's a messy, complex project, like writing a 50-page report or building a financial model. To judge this kind of work, we need a "rubric"—a detailed scoring sheet that tells us exactly what makes a piece of work excellent versus just okay. The problem is, most of these scoring sheets are made by guessing what the AI should do, or by looking at the AI's own answers, which is like asking a student to grade their own homework. We need a way to build these scoring sheets based on what real humans actually do in their jobs.

This paper introduces a new way to test financial AI agents called FinProBench, along with a clever method to build the scoring sheets, known as Role-Grounded Rubric Construction (RGRC). Think of RGRC as a detective that doesn't look at the suspects (the AI) to figure out the rules, but instead studies the crime scene (real, professional work documents) to see how the experts actually did it. The researchers gathered over 1,700 real documents—like investment reports and compliance checks—written by actual financial professionals. They used these real-world examples to teach the AI how to grade other AI work.

Here is the twist they discovered: It depends on the job. For common, everyday financial tasks that are already well-known (like writing a standard market report), simply asking the AI to "do a good job" (using a prompt) works almost as well as their fancy new method. However, for specialized, tricky roles where the rules are hidden or very specific (like a niche insurance actuary), the simple "ask nicely" approach fails miserably. In those cases, the RGRC method, which learns from real human deliverables, is a game-changer, catching 99.1% of the necessary standards compared to only 78% for the simple method.

When they put the AI agents to the test against real human experts using these new, high-quality scoring sheets, the humans still won on average, scoring 73.7 out of 100, while the best AI agents hovered around 70. But it wasn't a total blowout. The AI systems had their own superpowers: some were better at digging deep into the math, while others were great at formatting. The humans, however, were the most consistent all-rounders, nailing the structure, the data accuracy, and the compliance rules. The paper suggests that while AI is getting very good, it still needs to learn from the "tacit" (unspoken) secrets of real professionals to truly master complex jobs. The best part? Because they built a "master scoring sheet" for each job role, they could reuse it for many different tasks, saving a massive amount of time—about 6.7 times faster than making a new grading sheet from scratch for every single assignment.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →