← Latest papers
💬 NLP

Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable

This paper introduces the Capital Markets LLM Reliability Score (CM-LRS), a novel seven-dimensional evaluation framework designed to assess the "bankability" of large language model outputs in regulated financial workflows, revealing that while frontier models cluster closely in overall performance, significant gaps remain in retrieval and synthesis capabilities compared to open-weight baselines.

Original authors: Prerit Ahuja

Published 2026-07-28
📖 3 min read☕ Coffee break read

Original authors: Prerit Ahuja

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are in a giant library where a super-smart robot is hired to write reports for a bank. This robot, an Artificial Intelligence (AI) called a Large Language Model (or LLM), is incredibly talented at writing. It can sound like a seasoned expert, use perfect grammar, and tell a story that flows beautifully. But here is the catch: in the world of high-stakes finance, sounding good isn't enough. If the robot makes up a number, forgets to check its sources, or skips a step in a complex calculation, it could cost the bank millions or get them in trouble with regulators. The big question isn't "Can the robot write a smooth sentence?" but "Can the robot write a sentence that a human expert can trust with their career?" This paper dives into that exact problem, moving beyond simple tests of "did it get the answer right?" to a much deeper check of "is this answer safe to use?"

The authors introduce a new tool called the CM-LRS (Capital Markets LLM Reliability Score). Think of this as a "safety inspector" for AI reports. Instead of just giving a pass or fail grade, the inspector checks the report on seven different levels: Is the story true? Can you point to the exact page in the source document where the fact came from? Do the math numbers add up? Did the robot finish the whole job, or did it skip a step? Did it make up facts to fill gaps? Is the report actually useful for making a decision? And finally, can a human easily check the robot's work? The researchers tested four different AI models on five real-world banking tasks, like pulling numbers from bond contracts, finding similar past deals, and writing company profiles.

Here is what they found. First, the three most advanced, expensive "closed-source" AI models (from companies like Anthropic and OpenAI) are all performing at a very similar, high level of reliability. They are so close in score that it's hard to say one is clearly better than the others for these specific tasks. However, there is a clear gap between these top-tier models and one "open-source" model (Llama 3.3 70B). The open-source model scored significantly lower, especially on tasks that required digging through many documents or combining information from different sources. It was good at simple tasks like copying numbers from a single page, but it struggled when it had to find evidence or write a summary.

The most surprising discovery was about a specific part of the score called "Decision Usefulness." This measures whether a human banker would actually be willing to use the report without having to rewrite half of it. On one specific task (writing a company profile), the scores for this single factor varied wildly between the models, with a spread of 4.0 points on a 5-point scale. This suggests that while many AIs can write fluently, only the very best ones can produce work that is truly "bankable"—meaning it's ready to be used in the real world without needing a human to fix the mistakes. The paper concludes that for high-stakes jobs, we shouldn't just look at how well an AI talks; we need a strict checklist like the CM-LRS to ensure the math is right, the sources are real, and the work is safe to trust.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →