RT-SIM: A Simulation Framework for Evaluating Large Language Models for Trade Reconstruction and Reconciliation in Financial Operations
This paper introduces RT-SIM, a simulation framework demonstrating that frontier large language models can effectively reconstruct trades from trader communications for financial reconciliation, provided the dialogue is explicit or merely noisy, though performance significantly declines under ambiguous or semantically degraded conditions.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the financial world as a massive, high-speed game of trading cards, but instead of paper cards, people are swapping billions of dollars' worth of digital assets every second. In this chaotic arena, two things are absolutely critical: making sure you actually have the cards you think you have, and making sure your opponent isn't secretly holding a deck of fake ones. This is the job of "trade reconciliation," a fancy term for double-checking that the trade you made matches the trade the computer recorded. Usually, this is done by comparing two lists of numbers. But here's the twist: traders don't just type numbers into computers; they talk to each other. They chat, they joke, they argue, and they sometimes get confused. For decades, these conversations were ignored by the safety systems because they were too messy to read. Now, a new kind of computer brain called a Large Language Model (LLM) has arrived. Think of an LLM as a super-smart, tireless intern that can read thousands of messy conversations and try to figure out exactly what trade happened, even if the traders were talking over each other or using slang. The big question is: Can this super-intern actually do the job, or will it get lost in the noise?
This is exactly what Martin Higgins and Mackenzie Common set out to find out in their paper, "RT-SIM." They didn't just guess; they built a video game-like simulation called RT-SIM to test these AI brains in a controlled environment. Imagine they created a virtual trading floor where they could program traders to be perfect, slightly distracted, totally confused, or even speaking complete gibberish. Then, they asked different AI models to listen to these chats and write down the trades they heard.
The results were surprisingly clear. When the traders spoke clearly, the AI was amazing. Models like GPT-4o-mini and Claude Haiku 4.5 could reconstruct the trades with over 95% accuracy, even if the traders were throwing in a lot of irrelevant chatter about the weather or lunch. It turns out, the AI is great at ignoring the "noise" and finding the signal. However, the moment the traders started speaking in riddles, contradicting themselves, or getting confused, the AI's performance crashed. In the "confused" scenario, accuracy for some models dropped from nearly perfect to less than 10%. And when the traders spoke complete nonsense (gibberish), the AI didn't try to make things up; it simply admitted it found zero trades, which is actually a good thing because it means the AI isn't hallucinating fake deals out of thin air.
The researchers also looked at the cost and speed. They found that using these AI models is incredibly cheap—about $0.0002 per trade, which means you could process 10,000 trades for just $2.00. The speed was also fast enough to be useful in real-time. The main takeaway isn't that AI has solved all financial problems, but that it suggests a powerful new tool: if we can listen to the traders' conversations, we might be able to catch mistakes or even fraud much faster than we do today. However, the study warns that this tool only works if the conversations actually make sense. If the traders are too vague or contradictory, the AI can't help. So, while the technology is ready and affordable, it depends entirely on the quality of the human chatter it's trying to decode.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.