IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text
IndiaFinBench is the first publicly available benchmark designed to evaluate large language models on Indian financial regulatory text, featuring 406 expert-annotated question-answer pairs from SEBI and RBI documents that reveal significant performance gaps among models, particularly in numerical reasoning, while substantially outperforming non-specialist human baselines.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart robot librarian who has read almost every book in the world. You ask it a question about a library in New York, and it answers perfectly. But then you ask it a question about a very specific, old, and complicated library in India, full of unique rules and local slang. Suddenly, the robot stumbles. It doesn't know the local customs, and it gets confused by the specific way Indian laws are written.
This paper introduces IndiaFinBench, a new "test" designed to see how well these AI robots can handle the specific, tricky world of Indian financial regulations.
Here is the story of the paper, broken down into simple parts:
1. The Problem: The "Western Bias"
Until now, all the tests for AI in finance were like driving tests done only on American highways. They used US laws, US bank reports, and US news.
- The Reality: India has its own financial "highways" run by two main traffic cops: SEBI (for the stock market) and RBI (for the central bank).
- The Issue: Indian financial rules are messy. They are full of numbers hidden in paragraphs, rules that change over time (like a recipe that gets updated every year), and special words that only Indian experts know. Existing AI tests didn't check if the robots could handle this specific mess.
2. The Solution: The "Indian Driving Test"
The authors built IndiaFinBench, a brand-new exam with 406 questions.
- Where did the questions come from? They dug through 192 official documents from SEBI and RBI, ranging from 1992 to 2026.
- What kind of questions?
- The "What does this mean?" test: Can the AI read a rule and tell you exactly what a bank is allowed to do?
- The "Math" test: Can the AI do the math hidden in the text? (e.g., "If a bank has this much profit, how much dividend can they pay?")
- The "Lie Detector" test: Can the AI spot if two different rules contradict each other?
- The "Time Travel" test: Can the AI figure out which rule was active in 2015 versus 2023, given that rules often get updated?
3. The Exam Takers: 12 Different Robots
The researchers put 12 different AI models (the "robots") through this test. Some were huge giants (like Gemini 2.5 Flash and LLaMA-3.3-70B), and some were smaller, lighter models (like Gemma 4).
- The Rules: The robots had to answer using only the text provided in the question. They couldn't use their "memory" of the internet. This ensures they are actually reading and understanding, not just guessing.
4. The Results: Who Passed?
The results were surprising and revealed three distinct groups:
The Top Tier (The A-Students):
- Gemini 2.5 Flash came out on top with about 90% accuracy. It was the best at reading the rules and doing the math.
- Qwen3-32B and LLaMA-3.3-70B were right behind it, almost tied.
- The Surprise: Llama 4 Scout 17B (a much smaller robot) performed just as well as the massive 70B robot. It's like a compact car getting the same gas mileage as a massive truck. This suggests that how the robot is trained matters more than just how big it is.
The Middle Tier (The C-Students):
- Models like GPT-OSS and Mistral hovered around 75-79%. They were okay, but they made mistakes on the harder math and time-travel questions.
The Bottom Tier (The F-Student):
- Gemma 4 E4B (a very small model) scored around 70%. It struggled the most, especially with math and understanding the timeline of rules.
The Human Baseline:
Even a human who isn't a finance expert only scored 60%. This proves the test is actually hard! The best AI (Gemini) beat the average human by a huge margin, but the worst AI was barely better than a non-expert human.
5. Where Did the Robots Fail? (The Error Analysis)
Even the smartest robots had trouble spots:
- The "Time Travel" Problem: The hardest part for almost everyone was Temporal Reasoning. Indian rules often say, "This rule from 2010 is now replaced by the 2022 rule, unless you fall under the 2015 exception." The robots got confused about which rule applied when.
- The "Math" Problem: Numerical Reasoning was the biggest differentiator. The gap between the best and worst robot on math was huge (over 35%).
- The "Knowledge" Problem: Smaller robots often failed because they simply didn't know what words like "AIF" or "FEMA" meant. They were guessing.
6. Why This Matters
This paper is a wake-up call. It shows that:
- One size does not fit all: An AI trained on US laws isn't automatically good at Indian laws.
- Bigger isn't always better: A smaller, well-trained model can beat a giant one on specific local tasks.
- We need local tests: If we want AI to help Indian banks, lawyers, and regulators, we need to test them on Indian data, not just American data.
The Bottom Line:
The authors have released this test, the questions, and the answers to the public. They want everyone to keep testing new robots on Indian rules so we can build AI that truly understands the Indian financial world, not just the Western one. It's like giving the robots a local driver's license instead of just a US one.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.