From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate
This paper demonstrates that while decomposing European listed real estate analysis into specialized LLM agents significantly improves numerical task performance, enhancing integrative financial judgment requires targeted parameter adaptation via reinforcement learning (GRPO) rather than prompt-level decomposition alone.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a super-smart robot how to be a financial analyst. In the world of artificial intelligence, these robots are called Large Language Models (LLMs). Think of them as incredibly well-read students who have read almost every book in the library, but they sometimes get confused when asked to do specific, tricky math or to make a complex decision based on a mix of different rules. This paper lives in the corner of science where researchers try to figure out the best way to "train" these robots to handle money and business rules. The big question they are asking is: Is it better to ask the robot to do everything at once, or should we break the job down and ask it to act like a team of different specialists? It's a bit like asking, "Is it better to have one genius chef who tries to cook the whole banquet alone, or a kitchen with a dedicated baker, a dedicated grill master, and a dedicated sauce expert?"
The researchers in this study, Pardis Taghavi and Santosh Bhavani, decided to test this idea using a very specific and tricky subject: European real estate companies. These companies are like a group of neighbors who all live in the same building but follow different sets of rules depending on which country they are in. Some have to pay out most of their profits as dividends, while others don't have to pay anything at all. If you use the wrong rule for the wrong company, the math is wrong, and the investment advice is dangerous. The team wanted to see if splitting the analysis into "specialist" roles helped the robot get the numbers right, and if it helped the robot make the big, complicated judgments about whether a company is safe to invest in.
The Team of Specialists vs. The One-Person Show
To test their idea, the researchers built a system called "Larix." Imagine Larix as a manager who has a massive 16-step checklist for analyzing a real estate company. Instead of handing the whole checklist to one robot and saying, "Do this all by yourself," Larix tries a different approach. It breaks the checklist down into eight smaller teams, or "specialists." One team is the "Financials" expert, another is the "Asset Quality" expert, and another is the "Macro Overlay" expert.
The researchers ran a race with a very powerful robot (a "frontier" model) to see who would win. In the first race, the robot had to do everything alone (the "monolithic" approach). In the second race, the robot was given the full 16-step checklist but still had to do it all alone. In the third race, the robot was assigned to a specific specialist role, like a "Financials" expert, and only given the specific part of the checklist relevant to that job.
The results were surprising and very clear. When it came to the boring, mechanical math—like finding a specific number in a report or doing a simple calculation—the specialist approach was a huge winner. The robot acting as a specialist got the numerical tasks right 15.8 percentage points more often than the robot trying to do everything at once. It's as if the specialist robot put on noise-canceling headphones, ignored all the confusing extra instructions, and just focused on the math.
However, when it came to the "judgment" tasks—where the robot had to look at different pieces of evidence and decide if a company was risky or safe—the specialist approach didn't help at all. In fact, it sometimes made things worse. The robot acting as a specialist scored the same as the robot doing everything alone on these judgment tasks. It turns out that to make a good judgment, you actually need to see the whole picture, not just one narrow slice of it. By splitting the job up, the robot missed the connections between different pieces of information.
The Magic of "Training" the Brain
The researchers then asked a second question: If breaking the job into pieces doesn't help with the hard judgments, can we just "train" the robot's brain to get better at them? They took a smaller, cheaper robot (a 9-billion-parameter model called Qwen3.5-9B) and gave it a special kind of training called "Reinforcement Learning with Group Relative Policy Optimization" (GRPO).
Think of this training like a video game where the robot gets a high score every time it gets an answer right and a low score when it gets it wrong. Over time, the robot learns to adjust its internal "brain settings" to get more high scores. They didn't just ask it to memorize answers; they taught it to understand the rules of the game.
This training worked like magic for the judgment tasks. After the training, the smaller robot's score on the judgment tasks jumped up by 14.2 points. It didn't just get better at the specific companies it practiced on; it got better at companies it had never seen before. When tested on completely new firms, its score went up by 15.2 points overall, and on a specific "covenant stress" test (checking if a company can handle financial trouble), it improved by a massive 40.4 points.
The Big Takeaway
So, what did the paper actually find? It suggests that there is no single "best" way to use AI for financial analysis. It depends entirely on what you are asking the robot to do.
- For numbers and math: Breaking the job down into specialists is the way to go. It helps the robot focus and ignore distractions, leading to much more accurate calculations.
- For big-picture judgments: Breaking the job down doesn't help. Instead, you need to "train" the robot's brain directly so it learns how to connect the dots between different pieces of information.
The researchers were careful to say that this isn't a magic bullet that solves everything. They showed that just giving the robot a longer list of rules (the full 16-lens framework) didn't fix the problem; it was the specific specialization that helped the math, and the specific training that helped the judgment. They also noted that their results are based on a specific set of 19 companies and a specific type of real estate, so we can't be 100% sure it will work exactly the same way for every type of business in the world. But for the world of European real estate, the lesson is clear: use a team of specialists for the math, and use a well-trained generalist for the big decisions.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.