CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA
The paper introduces CLAIR-Fin, an adversarial nine-agent framework that enhances cross-modal financial question answering by decomposing queries into atomic claims for rigorous, type-aware verification and adaptive debate, thereby significantly improving faithfulness and reducing hallucinations compared to existing baselines.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery, but your clues are scattered across three very different types of notebooks: a handwritten diary full of stories, a spreadsheet packed with numbers, and a sketchbook filled with colorful charts. In the world of artificial intelligence, this is what happens when a computer tries to answer questions about complex financial reports. These reports are huge, often hundreds of pages long, mixing text, tables, and graphs together. The problem is that AI models, which are like super-smart but sometimes daydreaming students, often get confused when they have to read all these different formats at once. They might mix up a number from a table with a story from the text, or they might "hallucinate"—making up facts that sound real but aren't actually in the document. This is a big deal because in finance, a wrong number or a made-up story can lead to bad decisions. Scientists have been trying to build better ways for computers to check their own work, but most methods either trust all clues equally (even when a sketch is a bad clue for a specific number) or wait until the very end to check if the story makes sense, by which time it's too late to fix the mistakes.
Enter CLAIR-Fin, a new, clever system designed to act like a super-organized team of nine detectives working together to fact-check a financial report. Instead of letting one AI try to swallow the whole document at once, CLAIR-Fin breaks every question down into tiny, atomic "claims"—like individual puzzle pieces. For example, if the question is "How much money did the bank make last year?", the system doesn't just guess; it splits this into smaller claims like "The year was 2025" and "The profit was $139 million." Then, it sends these tiny claims to specialized agents. Some agents are experts at reading text, others at crunching numbers from tables, and others at interpreting charts.
Here is where the magic happens: CLAIR-Fin knows that not all clues are created equal. It uses a rule called "Asymmetric Evidence Authority." Think of it like a judge in a courtroom who knows that if you need an exact number, a spreadsheet cell is the gold standard, but if you need to know why something happened, a written story is the best evidence. The system trusts the spreadsheet for numbers and the story for reasons, rather than treating them all as the same. If the evidence is shaky or missing, the system doesn't force a guess; it simply admits it doesn't know and refuses to answer, which is a huge improvement over AI that just makes things up.
If a claim is tricky or the evidence is conflicting, the system kicks it into an "Adaptive Rebuttal Cycle." This is like a debate club where two AI lawyers—one arguing for the claim and one trying to poke holes in it—go back and forth. But here's the smart part: they only argue as long as necessary. If the claim is easy, they skip the debate. If it's hard, they argue more. This saves time and energy. Before the final answer is written, there is a "Chain-of-Custody" check, which is like a security guard making sure the evidence hasn't been swapped or tampered with during the hand-off between the researchers and the debaters. Finally, a "Judge-Auditor" gives a final verdict and calculates a "Hallucination Risk Index," which is like a weather forecast for how likely the answer is to be a sunny truth or a stormy lie.
The researchers tested this new system on a set of 500 questions based on real financial reports from the Bangladesh Bank. They found that CLAIR-Fin was much better at telling the truth than older methods. While a standard AI system got the facts right about 78% of the time, CLAIR-Fin raised that to nearly 89%. Even more impressively, when the evidence wasn't good enough to answer a question, the system chose to stay silent 5.4% of the time instead of forcing a wrong answer. This suggests that by breaking questions down, trusting the right type of evidence for the right job, and letting AI agents debate only when necessary, we can build AI that is much more reliable and honest when dealing with complex financial data.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.