BizFinBench.v2: Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation
The paper introduces BizFinBench.v2, a comprehensive bilingual benchmark utilizing authentic user data from Chinese and U.S. equity markets to evaluate large language models, revealing that current state-of-the-art models still fall short of practical business requirements despite DeepSeek-R1's superior performance.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you're trying to teach a super-smart robot how to be a stock trader. For years, we've been testing these robots in a giant, quiet classroom filled with fake practice exams. The robots ace the tests, getting A+ grades, so we assume they're ready to manage real money. But in the real world, the stock market isn't a quiet classroom; it's a chaotic, noisy, fast-moving rollercoaster where the rules change every second.
That's the big problem the authors of this paper, BizFinBench.v2, are tackling. They argue that the old "classroom tests" are lying to us. They are built on made-up data that doesn't match the messy reality of actual finance. So, the authors built a brand-new, real-world gym to test these robots.
The New Gym: Real Data, Real Stakes
Instead of fake questions, this new benchmark uses 28,860 actual questions that real people asked on real financial apps in both the Chinese and U.S. stock markets. It's like swapping a textbook quiz for a live, high-stakes game of poker where the chips are real.
The gym has two main zones:
- The Offline Zone: This is where the robots answer tough questions about things like "Why did this stock crash?" or "Analyze this company's report." There are 8 different types of tasks here, ranging from spotting weird data errors to predicting how a user feels about the market.
- The Online Zone: This is the scary part. Here, the robots have to make real-time decisions. They have to predict stock prices as they happen and manage a fake portfolio of money, buying and selling every hour based on live market data. It's like driving a car while blindfolded, but the car is moving at 100 mph and the road is changing every second.
The Results: The Robots Are Still Learning
When the authors put the world's smartest robots into this real-world gym, the results were... well, a bit of a shock.
Even the superstar robot, GPT-5, only got 61.5% of the answers right. That sounds okay for a school test, but in the real financial world, you need to be right 84.8% of the time just to be considered safe enough to use. The robots are still failing to meet the basic requirements of a real business.
Among all the robots tested, one named DeepSeek-R1 stood out, but with a catch. It was the only model that showed "superior investment efficacy" in the online investment game, actually making a profit of +47.34%. However, it was also a bit reckless, losing up to 53% of its money at one point (a "maximum drawdown"). It's like a race car driver who wins the race but crashes the car on the way to the podium. While it was the best of the bunch, it was still the only one that could even come close to the performance of mid-tier traditional trading strategies; the rest of the commercial models failed to beat the market.
Interestingly, the authors found that some robots that are famous for being "financial experts" actually performed worse than the general-purpose ones. It turns out that training a robot just on finance textbooks doesn't make it a better trader if it can't handle the chaos of the real market.
The "Thinking" Trap
The authors also tried a trick called Chain-of-Thought (CoT), where they told the robots to "think step-by-step" before answering. You'd think this would help, right? Wrong. For most robots, this actually made them worse. It was like telling a nervous student to over-analyze every single step of a math problem until they forgot the answer entirely. One robot, Claude-Sonnet-4, saw its score plummet from 39.5% down to 13.7% just because it was forced to think too hard!
The Five Big Mistakes
By looking at where the robots failed, the authors found five specific ways they mess up in the real world:
- Semantic Deviation: They understand the words but miss the meaning. They might think a tech company is the same as a semiconductor company just because they both use the word "tech," ignoring that they are totally different businesses.
- Logic Discontinuity: They lose the thread in long stories. If you ask them to trace a cause-and-effect chain over a long period, they get confused and drop the ball.
- Multivariate Analysis: They can't juggle too many balls at once. When they need to combine news, charts, and user feelings to make a decision, they get overwhelmed and pick the wrong factors.
- High-Precision Distortion: They are bad at math. In finance, a tiny decimal error can cost millions, and these robots often calculate numbers incorrectly.
- Time-Order Disorder: They get the timeline mixed up. They might think an event happened before it caused a reaction, or vice versa, completely flipping the story.
The Bottom Line
The paper suggests that while these AI models are getting smarter, they aren't ready to be your personal financial advisor just yet. The gap between their "test scores" and their "real-world performance" is huge. The authors aren't saying AI is useless; they are saying we need to stop testing them in fake classrooms and start testing them in the real, messy gym. Until they can consistently hit that 84.8% accuracy mark, we should be very careful letting them handle our money.
The good news? The authors are sharing their data and the "gym" they built so other scientists can help train the robots to be ready for the real deal. But for now, the robots are still just students, not masters.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.