Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents
This paper introduces \textsc{InvestLogicBench}, a large-scale benchmark of real-world investor decisions structured as Profile-Event-Reasoning-Decision-Outcome traces, to demonstrate that current financial LLMs exhibit polished but poorly grounded reasoning that is overlooked by traditional outcome-only evaluations, thereby advocating for a new process-native data interface for personalized consequential agents.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are walking into a giant library filled with every book ever written about money, stocks, and the economy. You have a super-smart robot friend who has read every single page. If you ask, "What is inflation?" or "Who founded Apple?", your robot can answer instantly and perfectly. This is how most computer programs called "Large Language Models" (LLMs) are tested today: they are like brilliant trivia champions who can recite facts and solve textbook math problems. But here is the twist: being a trivia champion doesn't mean you can play the game of life. In the real world, investing isn't about memorizing facts; it's about listening to a chaotic, noisy crowd of news, figuring out what actually matters, making a decision based on your own personality and risk tolerance, and then watching to see if you were right. It's the difference between knowing the rules of soccer and actually scoring a goal while the other team is trying to tackle you.
This is the exact problem a new paper tackles. The authors noticed that while these AI robots are amazing at answering questions on a test, they often fail miserably when asked to make real investment decisions in a live market. They created a new way to test AI, not by asking "What do you know?" but by asking "How do you think?" They built a special playground where they could watch how an AI connects the dots between a news event, its own reasoning, a specific action, and the final result. Their big discovery? The smartest AI robots, the ones that get the highest scores on general tests, are actually quite bad at this specific kind of thinking. They get confused by the noise, copy what everyone else is thinking, and often lose money, while some "less smart" models actually do better. The paper suggests that to build a truly useful financial robot, we need to stop testing its memory and start testing its ability to make a coherent plan in a messy, uncertain world.
The Great "Smart but Clueless" Paradox
The researchers set up a high-stakes game called a "live-trading arena." Imagine a video game where different AI bots are given $10,000 in fake money and told to trade stocks in the real U.S. market for seven weeks. They have to make their own choices every hour, reacting to breaking news just like a human trader would. The goal was to see if the AI's "general smarts" (how well it does on standard tests) matched its "trading smarts" (how much money it made).
The results were a bit shocking. The paper calls this the "Capabilities-Performance Paradox." It's like having a student who gets an A+ on every history exam but fails miserably at running a lemonade stand because they can't handle a sudden rainstorm. The AI models with the highest general scores, like GPT-5 and Claude-Sonnet-4.5, actually performed quite poorly. One of them, GPT-5, ended up with less money than it started with, losing about 1.0%. Another, Claude-Sonnet-4.5, barely made any profit at all. They both did worse than just sitting back and buying a simple index fund (the S&P 500), which is like the "do nothing" strategy.
In a twist of irony, the models that had lower scores on the general tests actually did much better in the market. DeepSeek-V3, for instance, turned its $10,000 into over $11,400, beating the market and the "smarter" models. This suggests that being good at answering questions doesn't automatically make you good at making decisions when the future is uncertain.
Why Do the "Smart" Bots Fail?
The authors dug deep to figure out why the high-scoring models were losing money. They found two main reasons, which they call "Consensus Bias" and "Reasoning Fragmentation."
Think of Consensus Bias like a sheep following the herd. When a big event happens, the "smart" models tend to just repeat what everyone else is already thinking because that's what they saw most often in their training data. They don't have the courage to look for a different angle or a hidden opportunity. They just say, "Oh, the market is going down, so I should sell," even if a human expert might see a chance to buy low.
Reasoning Fragmentation is like trying to solve a puzzle while someone is throwing confetti in your face. When the market gets noisy with lots of conflicting news, these models get overwhelmed. They can't connect the dots between a news story about a government meeting and a specific stock price. Instead of building a clear story, they get confused, jump to conclusions, or make decisions that contradict themselves. They might see a headline about a tech company and immediately buy it, ignoring the fact that the company just had a bad earnings report.
The New Test: INVESTLOGICBENCH2026
To fix this, the team built a new testing ground called INVESTLOGICBENCH2026. Instead of just asking the AI to answer a question, this benchmark looks at the whole chain of how an investment is made. They call it the P→E→R→D→O chain:
- P (Person): Who is the investor? What are their goals and fears?
- E (Event): What happened in the world? (e.g., "The Fed raised interest rates.")
- R (Reasoning): How does the investor connect the event to their goals? (e.g., "Higher rates mean borrowing is expensive, so tech stocks might drop.")
- D (Decision): What action do they take? (e.g., "Sell my tech stocks.")
- O (Outcome): Did it work? (e.g., "Yes, the stocks dropped, and I saved money.")
They gathered over 200,000 real investment decisions from 151 real-world experts, like famous fund managers and financial influencers. They broke down exactly how these humans thought, from the news they read to the trades they made. Then, they asked the AI models to do the same thing: look at the news, think like a specific type of investor, make a decision, and see if the logic holds up.
The Results: A Logic Gap
When they tested the AI models against this new standard, the results were clear. The models scored very low on "Reasoning Quality." On a scale of 1 to 5, the best-performing model (DeepSeek-V3) scored 2.8, while the worst (GPT-5) scored only 1.8. They struggled to tell the difference between important news and random noise. They often made decisions that didn't match their own reasoning or the investor's profile.
For example, in one test case, a human expert looked at a mix of news about a government shutdown, job reports, and interest rates, and correctly decided to buy "safe" assets. The AI models, however, got confused. Some focused only on the scary headlines, others ignored the most important parts, and some made contradictory arguments. One model even bought stocks that were completely unrelated to the news it was reading.
The paper concludes that while AI can generate profitable trades sometimes (perhaps by luck or by mimicking patterns), it lacks the deep, consistent logic that human experts use to navigate uncertainty. The "smart" models are failing not because they don't know the facts, but because they can't build a coherent story from the chaos of the real world.
What This Means for the Future
This paper doesn't say AI will never be good at investing. Instead, it says we need to change how we test and train them. We can't just keep giving them harder trivia quizzes. We need to teach them how to think like a strategist, how to filter out the noise, and how to stick to a plan even when things get scary. By using this new benchmark, researchers hope to build AI agents that are not just knowledgeable, but truly wise—capable of making personalized, logical decisions that actually work in the messy, unpredictable real world. The journey to a reliable financial AI is just beginning, and this new map is the first step in the right direction.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.