FinBalance: A Multi-Document Accounting Reconciliation Benchmark
The paper introduces FinBalance, a multi-document accounting reconciliation benchmark built from source documents across diverse industries and difficulty levels, which reveals that current large language models struggle to accurately reconcile source materials into consistent balance sheets, often exhibiting significant gaps between their reported financial statements and those derived from their own generated entries.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a financial mystery. You have a messy pile of clues: invoices, bank statements, contracts, and receipts. Your job isn't just to read them; it's to figure out exactly what happened, write down a formal report (a "journal entry") for every single transaction, and then add them all up to create a final "Balance Sheet" (a snapshot of the company's financial health).
If you find a clue that contradicts another (like an invoice saying you paid \100 but a bank statement saying you paid \120), you must flag that error immediately. You can't just ignore it or guess the answer.
This is exactly what the paper FinBalance is about. It introduces a new "exam" for Artificial Intelligence (AI) to see if it can do this real-world accounting detective work.
The Problem: AI is Good at Reading, Bad at Reconciling
The authors point out that most current AI tests are like giving a student a finished textbook chapter and asking, "What is the main idea?" The answer is already written there; the AI just has to find it.
But real accounting is harder. It's like giving the student a shoebox full of crumpled receipts, a bank statement, and a contract, and saying, "Figure out the math, write the official report, and tell me if anything doesn't add up."
The Solution: The FinBalance Benchmark
The researchers built a massive, synthetic "training ground" called FinBalance.
- The Setup: They created a computer program that acts like a strict accountant. It generates thousands of fake but realistic business scenarios across 8 different industries (like retail, healthcare, and tech).
- The Documents: For each scenario, it creates a bundle of documents with text inside them (simulating scanned PDFs). Some are real evidence, some are "distractors" (fake papers designed to trick you), and some contain deliberate contradictions.
- The Ground Truth: Because a computer generated the scenario, the AI knows the exact correct answer. It knows the right journal entries, the right final balance sheet, and exactly which documents support which numbers.
The Test: How Did the AI Do?
The researchers tested six of the smartest AI models available today on 710 of these accounting puzzles. The results were surprising and humbling:
The "Math vs. Report" Gap:
Imagine a student who can do the math perfectly in their head but writes the final answer on the test paper incorrectly.- The AI models were surprisingly good at figuring out the individual transactions (the "journal entries").
- However, when they tried to add those up and write the final "Balance Sheet," they often failed.
- The Stat: In four out of six models, there was a huge gap. They could reconstruct the correct financial picture from their own notes, but the report they actually submitted was wrong. It's like a chef cooking a perfect meal but serving it on the wrong plate with the wrong label.
The "Lost Receipt" Problem:
The biggest failure wasn't the math; it was linking the answer to the proof.- When the AI got a number right, it often couldn't point to the specific document that proved it.
- Even when the researchers told the AI, "You must cite your sources," it barely improved. The AI was struggling to find the needle in the haystack, not because it didn't want to, but because it couldn't distinguish the real evidence from the distractors.
The "Second Opinion" Trick:
The researchers tried a clever fix. After the AI submitted its answer, they ran the AI's own numbers through a strict calculator (a "ledger") and showed the AI the difference between what it said and what its own numbers implied.- Result: This "feedback" helped the AI fix its final report significantly. It's like a teacher saying, "You said the total is \100, but if you add your own list, it's \120. Check your math."
- The Catch: This fix had a downside. When the AI was trying to spot contradictions (errors in the documents), this feedback sometimes made it worse at spotting the errors, because it got too focused on fixing the math instead of realizing the documents were broken.
The Verdict
The paper concludes that while AI is getting better at understanding financial text, it is not yet ready to be a reliable accountant. It struggles to:
- Connect the dots between different documents.
- Aggregate (add up) its own findings into a consistent final report.
- Distinguish between real evidence and fake distractions.
The authors emphasize that this benchmark is a "stress test." If an AI can't solve these clean, computer-generated puzzles perfectly, it certainly can't be trusted with messy, real-world accounting where receipts are blurry, pages are missing, and the rules are even more complicated.
In short: AI is currently a great "reader" of financial documents but a poor "reconciler" of them. It can find the numbers, but it often loses the trail of evidence and fails to write the final report correctly.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.