FIND: Toward Multimodal Financial Reasoning and Question Answering for Indic Languages
This paper introduces FinVQA, a comprehensive multilingual benchmark for financial reasoning across six Indic languages, and proposes FIND, a framework combining supervised fine-tuning with constraint-aware decoding to enhance multimodal and numerical reasoning in high-stakes financial contexts.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: A New "Financial Gym" for AI
Imagine you are trying to teach a robot how to manage money. You wouldn't just ask it to read a book; you'd show it bank statements, stock charts, and receipts, and ask it to do the math.
The problem is that most AI robots are currently trained only in English and are terrible at looking at a chart and doing math at the same time. They often guess the answer based on the words they see, ignoring the numbers in the picture.
This paper introduces two things to fix that:
- FinVQA: A massive, difficult "exam" (dataset) for AI.
- FIND: A new "training method" (framework) to help AI pass that exam.
Both are specifically designed for Indic languages (like Hindi, Bengali, Tamil, etc.), which are spoken by millions in India but have been largely ignored by financial AI research.
Part 1: The Exam (FinVQA)
Think of FinVQA as a giant, multi-language test bank created by the authors.
- The Content: It contains nearly 19,000 questions. These aren't just simple "What is 2+2?" questions. They are complex financial puzzles involving:
- 14 different topics: From basic accounting and taxes to business economics and decision-making.
- 3 difficulty levels: Easy (like a warm-up), Moderate (a real workout), and Hard (the Olympic finals).
- Visuals: Many questions include images like graphs, tables, or receipts. The AI has to "read" the picture and the text together.
- The Languages: The test is available in English plus five major Indian languages: Hindi, Bengali, Marathi, Gujarati, and Tamil.
- The Goal: To see if an AI can actually reason through a financial problem in these languages, rather than just guessing.
How they built it:
The authors started with textbooks from India's National Council of Educational Research and Training (NCERT) and professional accounting courses. They took these English questions, translated them into the local languages, and even used AI to translate the text inside the images (like changing the labels on a chart from English to Hindi) so the visual clues matched the question.
Part 2: The Training Method (FIND)
Having a test is useless if the students (the AI models) don't know how to study for it. The authors propose FIND, a three-step training strategy.
Imagine you are teaching a student to take a math test.
Step 1: The "Zero-Shot" Attempt (The Guess):
You hand the student the test without any help. They try to solve it using only what they already know.- Result: They often get confused, write too much, mix up languages, or get the math wrong because they are rushing.
Step 2: The "Constrained Decoding" Rule (The Strict Proctor):
You tell the student: "You can think about the problem as much as you want, but when you write your final answer, you must only write the letter A, B, C, or D. No extra words."- Result: This stops the student from rambling or formatting their answer poorly. It forces them to be precise, but they might still get the math wrong.
Step 3: Supervised Fine-Tuning (The Tutor):
This is the big one. You sit with the student, show them the correct answers, and explain why they are right. You specifically train them to look at the charts and do the math correctly.- Result: The student learns to actually understand the financial concepts and do the calculations, not just guess.
The Finding: The paper shows that while the "Strict Proctor" (Step 2) helps with formatting, the "Tutor" (Step 3) is what actually makes the AI smart. When they combined the Tutor with the Strict Proctor, the AI performed significantly better, especially in the local Indian languages.
Part 3: What Happened When They Tested It?
The authors tested several different AI models (some small, some huge) on this new exam.
- Size Matters: Just like a human, a bigger brain (a larger AI model) generally does better. The biggest models (like the 32-billion parameter ones) got the highest scores, often exceeding 60-70% accuracy.
- The Language Gap: The AI did best in English, Hindi, and Bengali. It struggled more with Tamil, Marathi, and Gujarati, but the training method (FIND) helped close that gap significantly.
- The "Hallucination" Problem: Without the special training, smaller AI models would confidently explain their reasoning but get the final number wrong. It's like a student writing a perfect essay but getting the math calculation wrong. The FIND framework helped fix this, making the reasoning match the answer.
The Bottom Line
This paper says: "We built a hard, visual, multi-language financial test (FinVQA) and a way to train AI to pass it (FIND). We found that to make AI good at financial math in Indian languages, you need huge models, and you must train them specifically to look at charts and do the math, not just read the words."
Important Note from the Paper:
The authors warn that even with this training, AI is not perfect yet. It can still make small math errors or misunderstand complex charts. They emphasize that this tool is for research and evaluation only, not for giving real financial advice to people, because getting a number wrong in finance can have real-world consequences.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.