Can Open-Weight Models Compete on Financial Text Comprehension?
This paper evaluates updated Financial Touchstone benchmarks with twenty models, revealing that open-weight models like Kimi K2.6 can rival proprietary systems in financial text comprehension despite persistent information retrieval bottlenecks and inconsistent geopolitical refusal behaviors in Chinese models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where computers are learning to read, write, and reason like humans. This field, called Artificial Intelligence (AI), has recently exploded with new models that can chat, solve math problems, and even write code. But there's a catch: many of these super-smart computers are "black boxes." We know they work, but we don't know exactly how they were built or what secrets they hold inside their code. Recently, a new wave of these models has appeared from Chinese labs that are "open-weight," meaning their blueprints are public for anyone to study, unlike the secret, locked-up versions from big American tech companies. The big question on everyone's mind is: Can these open, public models actually do the hard, boring, but super-important work of reading financial reports and finding the truth, or are they just good at sounding smart?
This paper is like a massive, high-stakes test for twenty of these AI models. The researchers set up a "Financial Touchstone" challenge, which is basically a giant library of 495 real annual reports from companies all over the world, paired with 2,967 tricky questions about them. Think of it as giving a student a stack of 495 textbooks and asking, "How much profit did Company X make in 2022?" or "What are the three main business segments?" The goal was to see which AI could find the right answers without making things up (a problem called "hallucination") and without getting stuck because it couldn't find the right page in the book.
The results were a total plot twist. For a long time, people thought you needed a special "reasoning" brain or a secret, proprietary model to handle complex financial math. But this study found that open-weight models are catching up fast. In fact, the open-weight model called Kimi K2.6 came in third place overall, beating several famous, secret American models! The top performer was actually a model named Claude Opus 4.6, which got 88.4% of the answers right. However, there's a twist in the tale: while some models were great at finding the right answer, others were amazing at not making things up. Google's Gemini 2.5 Pro had the lowest rate of lying, with a hallucination rate of just 0.08%, but it wasn't the best at finding the answers in the first place.
The researchers also discovered that the biggest problem for all these AI models wasn't their brains; it was their eyes. Nearly half of all the mistakes (48.9%) happened because the computer simply couldn't find the right page in the massive report to read. It's like having a genius who can't find the book on the shelf. Another weird finding was that some Chinese models had a "safety filter" that would suddenly refuse to answer perfectly normal financial questions if the report mentioned certain political figures or events, like a student refusing to take a test because the topic reminded them of a rule they weren't supposed to break.
So, what's the bottom line? The paper suggests that open-weight models are now strong enough to compete with the best secret models in the world at understanding finance. They don't need to be "reasoning" super-computers to do a great job; sometimes, a simpler, open model works just as well. But, the study also warns that we still have a long way to go. The computers are getting better at reading, but they still struggle to find the right information in huge documents, and sometimes they get blocked by safety rules that are too sensitive. The author concludes that while we are close to having AI assistants that can handle real-world financial reports, we still need to fix how they search for information before we can fully trust them with our money.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.