FinSkillBench: Evaluating AI Agents and Domain Skills for Investment Management
The paper introduces FinSkillBench, a comprehensive benchmark for evaluating AI agents in investment management across portfolio construction, risk management, and fundamental analysis, demonstrating that access to curated procedural skills significantly enhances performance while self-generated skills offer minimal benefit.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the high-stakes world of investment management, success depends on more than just having a good idea about which stocks to buy. It requires a rigorous, step-by-step process of gathering precise data, performing complex calculations, and adhering to strict rules that govern how money can be moved. Imagine a financial analyst who must look at a portfolio of thirty different companies, calculate how they might react to a sudden market crash, and then reorganize the holdings to stay within legal limits, all while using only the information available on a specific day in the past. This is the daily reality for human professionals, but it is a new frontier for artificial intelligence. While modern AI systems are excellent at writing text and answering general questions, they often struggle when asked to perform these exact, disciplined tasks. They might understand the concept of risk, but they frequently fail to execute the specific mathematical steps required to measure it accurately or to follow the detailed instructions needed to rebalance a portfolio without error.
A team of researchers has created a new way to test whether AI agents can truly handle this kind of work. They built a specialized testing ground called FinSkillBench, which presents AI systems with 2,603 distinct investment scenarios. These scenarios cover three main areas: building a portfolio of stocks, managing the risks associated with those stocks, and analyzing the financial health of individual companies. The researchers did not just ask the AI to guess the answer; they gave it a specific date, a set of raw data, and a clear goal, then checked the result against a known, correct solution generated by standard financial software. The goal was to see if the AI could act like a professional analyst, retrieving the right data at the right time and using the right tools to get the job done.
The study tested three different ways of helping the AI. In the first scenario, the AI had to solve the problems entirely on its own, with no help. In the second, the researchers provided the AI with a "curated" set of tools and written guides. These were not just random documents; they were carefully prepared instructions and computer scripts that showed the AI exactly how to perform the necessary calculations and how to use the available tools correctly. In the third scenario, the AI was asked to write its own instructions and tools for itself before attempting the task, essentially trying to teach itself how to do the job on the fly.
The results were clear and decisive. When the AI was given the pre-made, human-written guides and tools, its performance improved dramatically. Across all the tests, the average score for these AI agents jumped from roughly 37 percent to 53 percent. The improvement was most significant in the tasks that required heavy number-crunching, such as building an optimal portfolio or identifying hidden risks. In these areas, the AI with the curated tools was able to solve problems that it previously could not solve at all. The researchers found that having access to these reliable, step-by-step procedures was just as important as the choice of the AI model itself. A powerful AI model without the right tools often failed, while a capable model with the right tools succeeded.
In contrast, the idea of the AI teaching itself on the spot did not work. When the agents were forced to write their own guides and tools before solving the problems, their performance barely changed. They spent a lot of time and computing power writing these instructions, but the instructions they created were often flawed or incomplete. The study showed that in a single attempt, an AI cannot reliably invent the complex, precise procedures needed for financial analysis. It is not enough for the AI to know how to write code; it needs to be given code that has already been tested and verified. The researchers also ran the same tests using a different AI system to make sure their findings were not just a fluke of one specific setup. The second system produced the same pattern: pre-made tools helped, and self-made tools did not.
This work suggests that for AI to be useful in serious fields like finance, we cannot rely on the models to figure out the rules of the game as they go along. Instead, we must build systems where the AI has access to trusted, human-verified methods for doing the hard work. The study concludes that the future of AI in investment management lies not in smarter models alone, but in better-designed systems that combine those models with reliable, reusable skills. The researchers have made their tests and tools available to others, hoping to help the field move forward by focusing on how AI agents can be equipped with the right resources to solve real-world problems.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.