← Latest papers
💻 computer science

DiligenceProv: A Truth-Ledger Benchmark for Verifiable Financial Due-Diligence Answers from Large Language Models

This paper introduces DiligenceProv, a novel benchmark and construction methodology using synthetic deal rooms with registered imperfections to evaluate and improve the verifiability, evidence retrieval, and definitional accuracy of large language models in high-stakes financial due diligence.

Original authors: Yunguo Yu

Published 2026-09-14
📖 6 min read🧠 Deep dive

Original authors: Yunguo Yu

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

When a company is bought or sold, the process relies on a rigorous check of its financial health, a stage known as due diligence. Advisors spend weeks poring over a "deal room," a collection of documents ranging from audited financial statements and earnings reports to draft contracts and management presentations. Their job is to answer critical questions: How much does the business truly earn? What risks do its largest customers carry? In this high-stakes environment, an answer is only as good as the proof behind it. A reviewer must be able to verify not just the final number, but the specific definition used to calculate it and the exact document that supports the claim. If an answer looks confident but is wrong, it can mislead an investment decision more dangerously than no answer at all.

Artificial intelligence, specifically large language models, is beginning to enter this field, promising to speed up the reading and analysis of these complex documents. However, a significant obstacle remains. Current AI systems can read text fluently, but they struggle with the specific kind of verification required in finance. They might produce a plausible number that is actually incorrect, or they might cite a document that does not support their conclusion. The core challenge is not whether the AI can read the words, but whether it can navigate a messy collection of conflicting information to find the single truth that a human expert would accept.

To address this, researchers have created a new testing ground called DiligenceProv. This is not a standard test based on clean, public financial reports where the answers are already reconciled and agreed upon. Instead, the researchers built a fictional company and generated a realistic set of thirty documents, including financial statements, debt schedules, and management presentations. They deliberately introduced the kinds of messiness found in real-world deals: the same earnings measure defined in three different ways, outdated numbers that look current, and questions that the documents simply cannot answer. The goal was to see if an AI system could provide answers that a human reviewer could verify, checking the number, the definition, and the evidence all at once.

The researchers found that simply making the documents look realistic was not enough to challenge the most advanced AI systems. In an initial test with a realistic but perfectly consistent set of documents, a top-tier commercial model answered every question correctly, even when the questions were tricky. The system had no trouble because the documents did not contain the specific kinds of traps that exist in real negotiations. The researchers realized that to truly test these systems, they had to engineer the difficulty. They created a "truth ledger," a master record of every fact, and then wrote the documents to contain registered imperfections. These included conflicts between different documents, realistic distractions that looked like the right answer but were wrong, and silent gaps where information was missing but never explicitly stated as missing.

When they tested the AI models against this engineered mess, the results changed dramatically. The most advanced models, which had previously scored perfectly, began to fail when they had to navigate these registered imperfections. More surprisingly, the researchers discovered that for many of the mid-range models, giving the AI the entire set of thirty documents actually made it perform worse than giving it just six carefully selected pages. When the AI had access to the full room, it often got distracted by outdated numbers or fabricated citations that appeared in the extra text. It would grab the wrong version of a number or invent a document that did not exist. However, when the researchers provided only the most relevant passages, the AI performed significantly better. In this case, the retrieval process acted as a protective filter, shielding the model from the noise and confusion of the full document set.

The study also revealed that the hardest part for these AI systems was not finding the evidence, but understanding how to use it. Even when the researchers gave the models the exact correct documents and the right numbers, the models still made mistakes. They often failed to apply the correct definition of a financial metric or could not distinguish between a draft version of a contract and the final one. The errors were not due to a lack of information, but a lack of discipline in applying the rules of the deal. The researchers categorized these failures into specific types, such as "confident-wrong answers" where the AI states a false fact with certainty, or "distractor capture" where it latches onto a misleading number.

A particularly troubling finding was the AI's inability to admit when it did not know something. In the real world, if a document does not contain a specific forecast, the correct answer is to say that the information is missing. In the tests, the AI models rarely refused to answer. Instead, when faced with missing data, they often tried to guess or extrapolate, creating a fabricated number that looked structured and professional. This is a dangerous behavior in finance, where a confident-sounding wrong number can be more harmful than a clear admission of missing information.

The researchers concluded that the difficulty in financial due diligence is not just about the complexity of the text, but about the way evidence is accessed and managed. For the strongest AI systems, having access to the full context of all documents allowed them to reach near-perfect scores, suggesting that their main limitation is simply having the right information available. For other models, having too much information was a liability. The study suggests that for these systems to be trusted with real deals, they need workflows that prioritize verification and definition over simple retrieval. The benchmark itself is now available for other researchers to use, providing a way to measure and improve these systems before they are trusted with real money. The work demonstrates that for high-stakes financial tasks, the path to reliability lies not in making the AI smarter, but in building systems that can verify their own answers against a strict set of rules and evidence.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →