Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation
Conv-FinRe introduces a conversational and longitudinal benchmark for stock recommendation that evaluates large language models by distinguishing between imitating noisy user behavior and generating utility-grounded rankings aligned with long-term investor goals, revealing a persistent tension between rational decision quality and behavioral alignment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a financial advisor. You have two ways to judge if they are good:
- The "Copycat" Test: Did they tell you to buy exactly what you bought last time?
- The "Wisdom" Test: Did they give you advice that would actually make you richer and safer in the long run, even if it meant telling you not to do what you wanted to do?
Most current AI tests for financial advice only use the Copycat Test. They ask, "Did the AI guess what the user clicked?" The problem is, humans are messy. Sometimes we panic and sell stocks when the market dips, or we get greedy and buy a risky stock because it's trending. If an AI just copies those messy, emotional moves, it's not being a good advisor; it's just being a bad mirror.
Conv-FinRe is a new, smarter test designed to fix this. Here is how it works, broken down with simple analogies:
1. The Setup: A Long-Term Relationship, Not a One-Time Chat
Imagine you are hiring a personal trainer.
- Old Tests: You ask the trainer, "What workout did you do yesterday?" and they just repeat your exact routine.
- Conv-FinRe: This benchmark simulates a 30-day relationship. It starts with an interview (the "onboarding") where the AI learns your goals, your fears, and your budget. Then, day by day, the AI talks to you as the market changes, giving advice and seeing how you react. It's not a one-off question; it's a long conversation.
2. The "Four Eyes" of Judgment
This is the most creative part. When the AI gives a recommendation, the test doesn't just check one answer key. It checks the AI's advice against four different "expert eyes" to see what the AI is actually thinking:
- 👁️ The "User Eye" (What you actually did): Did the AI copy your past messy choices? (This is the "Copycat" score).
- 🧠 The "Rational Eye" (What is mathematically best): Based on your risk tolerance, what should a perfect, calm investor do? (This is the "Wisdom" score).
- 📈 The "Hype Eye" (What the market is doing right now): Did the AI just follow the crowd and buy the hottest stock of the week?
- 🛡️ The "Safety Eye" (What keeps you safe): Did the AI prioritize avoiding big losses over chasing big gains?
The Analogy: Imagine a student taking a test.
- If they just copy their friend's answers, they pass the "User Eye."
- If they solve the problem using perfect logic, they pass the "Rational Eye."
- If they just guess the answer that's trending on social media, they pass the "Hype Eye."
- Conv-FinRe wants to know: Is the AI a smart logician, a mindless copycat, or a trend-chaser?
3. The Secret Sauce: "Reverse Engineering" Your Brain
How does the test know what the "Rational" answer is if the user made a "messy" choice?
The researchers used a clever trick called Inverse Optimization.
- Think of it like this: You see a chess player make a weird move. You don't know their strategy. But by watching many of their moves over time, you can mathematically "reverse engineer" their brain to figure out: "Ah, this player is actually very risk-averse; they avoid losing pieces even if it means winning fewer points."
- The benchmark does this with the user's data. It figures out the user's true hidden personality (e.g., "I hate losing money more than I love making it") and uses that as the "Gold Standard" for what good advice looks like.
4. The Big Discovery: The "Tug-of-War"
When they tested top AI models (like GPT-4, Llama, etc.) on this new benchmark, they found a funny and important tension:
- The "Smart" AIs: Some models were great at the Rational Eye. They gave mathematically perfect advice that aligned with the user's long-term goals. However, they often failed the User Eye. They told the user to do something the user wouldn't actually do, so the user felt misunderstood.
- The "Empathetic" AIs: Other models (especially those trained specifically on finance) were great at the User Eye. They mimicked the user's messy, emotional choices perfectly. However, they failed the Rational Eye. They were so busy copying the user's bad habits that they gave terrible financial advice.
The Bottom Line
Conv-FinRe teaches us that a truly great financial AI shouldn't just be a yes-man who copies your mistakes, nor should it be a robot who ignores your feelings.
It needs to be a wise guide: someone who understands your personality, respects your fears, but gently steers you away from the emotional traps that hurt your wallet. This new benchmark is the first tool that can actually measure if an AI is doing that balancing act.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.