FinDeepIndicator: Benchmarking Deep Research Agents in End-to-End Financial Indicator Construction
This paper introduces FinDeepIndicator, the first benchmark designed to evaluate Deep Research agents across the four end-to-end stages of financial indicator construction, revealing that while these agents outperform standard search-equipped LLMs, they still struggle with reliability in data retrieval and numerical execution within realistic financial analysis scenarios.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery, but instead of just asking a witness "What did you see?", you have to do the whole job yourself. You have to figure out what question to ask, find the witness in a crowded city, listen to their story, write it down, do the math to see if the story adds up, and finally, tell the judge the verdict. This is the world of "Deep Research" agents. These are super-smart computer programs designed to act like human researchers: they don't just answer questions; they go out, hunt for information, piece it together, and solve complex problems from scratch. But here's the catch: just because a computer can talk like a human doesn't mean it can act like a human. It might know the definition of a word perfectly but fail to find the right book in the library or mess up the addition when counting the pages.
In the world of finance, this is a big deal. Financial indicators are like the scorecards of the economy—numbers that tell us if a company is healthy, if the stock market is nervous, or if the country is growing. To get these numbers, you can't just guess; you have to dig through years of financial reports, find the exact numbers for specific companies, and crunch the data. Scientists have been building these "Deep Research" agents to do this work, hoping they can replace or help human analysts. But until now, no one had really tested if these agents could actually handle the whole messy process from start to finish, or if they were just good at sounding smart while making silly mistakes in the middle.
Enter FinDeepIndicator, a new "exam" created by a team of researchers to test these digital detectives. Think of this benchmark as a giant obstacle course designed specifically for financial agents. Instead of just asking, "What is the profit of Company X?", the exam forces the agent to do four distinct things: first, figure out the exact math formula needed (like knowing that "Profit Margin" means dividing profit by sales); second, go out and find the raw numbers from the internet; third, actually do the math to get the intermediate results; and finally, give the correct final answer. The researchers built this course using 3,350 different questions covering 234 different types of financial scores, pulling data from 800 real companies in both the U.S. and China over a 10-year period.
When they ran the agents through this obstacle course, the results were a mix of "not bad" and "oh no." The agents were surprisingly good at the first step: understanding the rules. They could correctly identify the formulas more than 70% of the time, showing they really do understand what the financial terms mean. However, the moment they had to go out and find the data, things fell apart. The accuracy dropped like a stone. The agents struggled to find the right numbers, often grabbing the wrong year, the wrong company, or missing a crucial piece of information entirely. This "data collection" step turned out to be the biggest bottleneck, causing the biggest drop in performance.
Even the smartest agents, which could plan their search and use tools better than a simple search engine, only got the final answer right about 40% of the time. The researchers found that the agents were particularly bad at "macroeconomic" indicators (like inflation or labor rates), where the data is messy and comes from different sources, and they struggled even more with "technical" indicators that require complex, rolling calculations. Interestingly, the agents performed slightly better on U.S. market data than on Chinese market data, suggesting that the way data is organized and reported in different countries makes a huge difference.
The bottom line is that while these AI agents are great at knowing the theory of finance, they are still clumsy at the practice of digging up the facts and doing the math. The paper suggests that if we want these agents to be truly useful in the real world, we need to stop focusing just on how well they can chat and start fixing their ability to find reliable data and process it without making mistakes. Until then, they are like a brilliant detective who knows the law perfectly but keeps getting lost on the way to the crime scene.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.