FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs
This paper introduces FinAuditing, a large-scale, taxonomy-aligned benchmark derived from real XBRL filings that evaluates the capabilities of 13 state-of-the-art LLMs across three financial auditing tasks, revealing significant gaps in their ability to perform structure-aware semantic matching, relationship extraction, and mathematical reasoning.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a massive financial mystery. The suspect is a giant corporation, and the evidence is a stack of thousands of pages of financial reports. But here's the catch: these aren't just normal pages of text. They are like a giant, interconnected puzzle made of Lego bricks, where every single number, word, and relationship has to fit perfectly according to a strict rulebook called US-GAAP (the "constitution" of accounting).
This paper introduces a new tool called FINAUDITING, which is essentially a "final exam" for Artificial Intelligence (AI) to see if it's smart enough to be a financial detective.
Here is the breakdown of the problem and the solution, using simple analogies:
1. The Problem: The "Invisible" Mistakes
Imagine a company says, "We have $100 million in cash!" in their summary. But when you look at the detailed receipts in the back of the report, the math only adds up to $95 million.
In the old days, a human auditor would have to manually check every single page to find this $5 million gap. It's like trying to find a specific typo in a library of a million books. It's slow, boring, and humans get tired.
Now, we have Large Language Models (LLMs)—super-smart AIs that can read and understand text. But can they do this specific job?
- The Issue: Most AIs are great at writing poems or chatting. But financial auditing isn't just about reading words; it's about checking structure, math, and rules.
- The Gap: Existing tests for AI are like asking, "Can you write a sentence about money?" FINAUDITING asks, "Can you find the $5 million lie hidden in a 30,000-word legal document that follows a strict 1,000-page rulebook?"
2. The Solution: FINAUDITING (The "Financial Gym")
The authors built a massive, realistic training ground (a benchmark) using real-world data from the SEC (the US financial police). They created three specific "workouts" to test the AI's muscles:
Workout A: The "Name Tag" Test (Financial Semantic Matching)
- The Analogy: Imagine a library where every book must have a specific label on its spine. If a book is about "Apples," it must be labeled "Fruit," not "Vegetable."
- The Task: The AI has to look at a financial report and check if the company used the right "official labels" (tags) for their numbers. Did they call "Cash" what the rulebook says it should be called?
- The Result: The AIs struggled. They often picked the "closest sounding" word instead of the correct legal word. It's like calling a "Spatula" a "Spoon" because they both look like kitchen tools, even though they aren't the same.
Workout B: The "Family Tree" Test (Financial Relationship Extraction)
- The Analogy: Think of a family tree. A "Father" must be the parent of a "Son." You can't have a "Grandson" listed as the parent of a "Father."
- The Task: Financial reports have hierarchies. "Total Assets" must be the parent of "Current Assets." The AI has to check if the company messed up the family tree. Did they put a "Child" above a "Parent"?
- The Result: This was very hard. The AIs got confused by the complex structure. They could read the words, but they couldn't understand the relationships between them.
Workout C: The "Math Check" Test (Financial Mathematical Reasoning)
- The Analogy: You have a receipt that says "Total: $50." But the items listed are 20 + $15. The math doesn't add up.
- The Task: The AI has to look at the numbers, find the hidden math rules (like "Total = Sum of Parts"), and calculate if the company's numbers are lying.
- The Result: This was the hardest part. Even the smartest AIs failed to do the math correctly. They could find the numbers, but they couldn't do the arithmetic or follow the chain of logic to prove the lie.
3. The Big Reveal: AI Isn't Ready Yet
The authors tested 13 of the smartest AI models in the world (including giants like GPT-4o and specialized financial AIs).
- The Verdict: The results were disappointing. Even the "super-smart" models got most of the questions wrong.
- Why? These AIs are like brilliant readers who are terrible accountants. They can understand the story, but they don't understand the strict rules, the math, or the deep structure required for real auditing. They rely on "guessing" based on patterns rather than "reasoning" based on facts.
4. Why This Matters
This paper is a wake-up call.
- For the Public: It means we can't just trust an AI to audit our money yet. If we let them loose, they might miss huge scandals because they don't understand the "rules of the game."
- For the Future: The authors released this "exam" for free. Now, other scientists can use it to build better AIs that actually understand structure and math, not just words.
In a nutshell:
FINAUDITING is a reality check. It shows that while AI is amazing at writing and chatting, it is currently not smart enough to be a financial auditor. It needs to learn how to follow strict rules, check its own math, and understand complex structures before it can be trusted with our money.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.