← Latest papers
🤖 AI

Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows

This paper introduces FinWorkBench (Finch), a comprehensive benchmark derived from authentic enterprise data that evaluates frontier AI agents on complex, multimodal finance and accounting workflows, revealing significant performance gaps where even advanced models like GPT 5.1 Pro achieve only a 38.4% success rate.

Original authors: Haoyu Dong, Pengkun Zhang, Yan Gao, Xuanyu Dong, Yilin Cheng, Mingzhe Lu, Zikun Zhu, Adina Yakefu, Shuxin Zheng

Published 2026-04-07
📖 5 min read🧠 Deep dive

Original authors: Haoyu Dong, Pengkun Zhang, Yan Gao, Xuanyu Dong, Yilin Cheng, Mingzhe Lu, Zikun Zhu, Adina Yakefu, Shuxin Zheng

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a super-smart robot how to do your job as a financial analyst. You might think, "No problem! Just tell it to 'add up the numbers' or 'make a chart.'"

But in the real world of finance, it's not that simple. It's less like solving a math problem on a clean whiteboard and more like trying to fix a leaky roof while standing on a wobbly ladder, in the rain, holding a map that's been crumpled and taped together, and the map is written in three different languages.

That is exactly what the paper FINCH is about.

Here is the breakdown of the paper in plain English, using some everyday analogies.

1. The Problem: The "Messy Kitchen" of Finance

Most AI tests today are like asking a robot to bake a cake using a perfect, pre-measured recipe in a pristine kitchen. The ingredients are in neat bowls, the instructions are clear, and the oven is set to the right temperature.

But real finance work is a messy kitchen.

  • The Ingredients are scattered: Data isn't in one neat list. It's in 50 different Excel files, some PDFs, some emails, and some charts.
  • The Recipe is broken: The instructions are hidden inside old emails ("Hey, can you update the 2002 budget?") or buried in the history of how a spreadsheet changed over time.
  • The tools are weird: The "tables" aren't perfect grids. They have merged cells, weird fonts, charts embedded in the middle of text, and formulas that look like secret codes.
  • The job is long: You don't just "add numbers." You have to find the data, clean it, check if it makes sense, calculate new projections, draw a graph, and write a report. It's a long chain of steps where one mistake ruins the whole thing.

2. The Solution: FINCH (The "Real-World Gym")

The authors created a new benchmark called FINCH (Finance & Accounting Benchmark). Think of this as a gym for AI robots, but instead of lifting weights, they have to navigate a messy office.

  • Where did they get the data? They didn't make up fake problems. They dug into real, historical archives (like the famous Enron emails and real bank documents) from 2000 to 2025. They found 172 real-life tasks that actual humans had to do.
  • How did they build it? They used a mix of AI and human experts. Imagine a team of detectives looking at an old email thread and a stack of spreadsheets, trying to figure out: "What was the human trying to do here? What steps did they take? What was the final result?" Then, they wrote a "test" based on that story.
  • The Scale: This isn't a small test. It involves 1,710 spreadsheets and 27 million cells. It's a massive, chaotic library of financial data.

3. The Test: Can the Robots Do It?

The authors took the smartest AI robots available today (like GPT-5, Claude, and Gemini) and threw them into this messy kitchen.

The Results were shocking:
Even the "super-brains" failed most of the time.

  • The Pass Rate: The best AI only got about 38% of the tasks right. That means they failed more than 6 out of 10 times.
  • The Time: It took the AI an average of 17 minutes to try to solve one task, and they still got it wrong.
  • The "Long Chain" Problem: If a task had just one or two steps, the AI did okay. But if the task had three or more steps (like: Find data -> Clean it -> Calculate -> Draw Chart), the AI's success rate plummeted. It's like a robot that can tie one shoe but forgets how to tie the second one, or gets confused by the third.

4. Why Do They Fail? (The "Gotchas")

The paper found five main reasons why the robots are struggling:

  1. The "Where is it?" Problem: The data is spread across 10 different files. The AI often looks in the wrong file or grabs the wrong row of numbers. It's like looking for a specific ingredient in a pantry with 1,000 unlabeled jars.
  2. The "Secret Code" Problem: In finance, the formula in a cell is often more important than the number you see. The AI sees the number "100" but misses the hidden logic that says, "This 100 is actually a 55-day payment schedule." It ignores the instructions behind the curtain.
  3. The "Messy Table" Problem: Real spreadsheets are ugly. They have merged cells, blank rows, and weird headers. The AI gets confused by the layout and thinks a header is a data point, or misses a column entirely.
  4. The "Multimodal" Problem: Sometimes the data is in a PDF, sometimes in a picture of a chart, and sometimes in an email. The AI struggles to combine these different formats into one coherent picture.
  5. The "Memory" Problem: Because the tasks are so long, the AI forgets what it did in step 1 by the time it gets to step 10. It loses the thread of the story.

5. The Big Takeaway

The paper concludes that while AI is amazing at writing poems or coding simple scripts, it is not yet ready to replace a human financial analyst.

Real finance work is too messy, too long, and too full of hidden context for current robots to handle reliably. The AI is like a brilliant student who can ace a multiple-choice quiz but freezes when asked to fix a real-world plumbing leak.

The Future:
The authors hope FINCH will become the standard "driving test" for financial AI. Until AI can pass this test, we shouldn't trust it to manage our money, our budgets, or our taxes without a human looking over its shoulder.

In short: FINCH is a reality check. It shows us that the "messy middle" of real-world work is still very hard for AI to navigate.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →