← Latest papers
💬 NLP

FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents

This paper introduces FinPersona-Bench, a simulation benchmark that reveals how autonomous financial agents suffer from Mandate Salience Decay over time, demonstrating that while periodic mandate re-grounding can stabilize conservative agents, it may inadvertently worsen the behavior of aggressive ones depending on market conditions.

Original authors: Muhammad Usman Safder (Steve), Ayesha Gull (Steve), Rania Elbadry (Steve), Fan Zhang (Steve), Yankai Chen (Steve), Xueqing Peng (Steve), Xue (Steve), Liu, Preslav Nakov, Zhuohan Xie

Published 2026-07-01
📖 5 min read🧠 Deep dive

Original authors: Muhammad Usman Safder (Steve), Ayesha Gull (Steve), Rania Elbadry (Steve), Fan Zhang (Steve), Yankai Chen (Steve), Xueqing Peng (Steve), Xue (Steve), Liu, Preslav Nakov, Zhuohan Xie

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you hire a personal financial advisor. Before they start, you give them a very clear, written rulebook: "Your only job is to protect our money. Never take big risks. If things get scary, hold tight and don't sell."

You expect them to follow this rule forever. But this paper, FinPersona-Bench, discovered something surprising: the longer the advisor works, the more they forget their own rulebook.

Even if the advisor is a super-smart AI (a Large Language Model), as they see more daily market news, price charts, and market noise, their original instructions start to fade into the background. They begin to act more like the market around them and less like the person you hired them to be.

Here is a breakdown of the paper's findings using simple analogies:

1. The Problem: "The Fading Tattoo"

The authors call this phenomenon Mandate Salience Decay (MSD).

Think of your rulebook like a tattoo on the advisor's arm.

  • At the start: The tattoo is fresh, bright, and impossible to ignore. The advisor follows it perfectly.
  • After 200 days: The tattoo hasn't disappeared, but it's faded. The advisor is now surrounded by a loud, chaotic crowd (the market). The crowd's shouting is so loud that the faint tattoo is drowned out. The advisor starts listening to the crowd instead of their own skin.

The paper proves that even if the AI is smart enough to make a profitable trade, it might violate its own core rule (like "don't panic sell") just because the market context has overwhelmed its original instructions.

2. The Experiment: A "Fake" Market

To prove this, the researchers built a video game version of the stock market.

  • The Secret: In this game, there is a "True Value" for every stock (like the real weight of a gold bar), but the AI agents cannot see it. They only see the "Market Price" (the sticker price on the bar), which can be manipulated to be higher or lower than the truth.
  • The Test: They gave different AI agents different "personalities" (like a cautious "Guardian" who hates risk, or an aggressive "Commander" who loves growth).
  • The Scenarios: They put these agents through three types of games:
    1. The Boring Flat Market: Nothing happens. The test is: Will the agent stay calm and do nothing, or will it get bored and start trading for no reason?
    2. The Crash: Prices plummet. The test is: Will the agent panic and sell everything, or will it stick to its plan?
    3. The Bubble: Prices go crazy high. The test is: Will the agent get greedy and buy the bubble, or stay rational?

3. The Results: "Memory" Helps, But Not Always

The researchers tried a fix: Re-grounding.
Instead of just telling the agent the rules once at the start, they reminded the agent of the rules every single day (like a daily pep talk).

What happened?

  • The Good News: For the Cautious Guardian agents, the daily reminders worked perfectly. They stayed calm in boring markets and didn't panic during crashes. The reminders acted like a safety anchor.
  • The Bad News: For the Aggressive Commander agents, the daily reminders actually made things worse in boring markets. Because the market was quiet, the aggressive agent was told every day to "Go get those big wins!" This made them trade too much and lose money, even though the market had no reason to move.

The Analogy:
Imagine a Guardian is a turtle. If you remind a turtle every day to "stay in your shell," it stays safe.
Imagine a Commander is a race car. If you remind a race car every day to "go fast" while it's sitting in a parking lot (a boring market), it will crash into the walls. The reminder didn't help; it forced the car to act against the reality of the situation.

4. Key Takeaways

  • Time is the enemy: The longer an AI agent runs, the more its original instructions fade. The gap between what it should do and what it actually does grows bigger every day.
  • One size does not fit all: Reminding an agent of its rules helps some personalities but hurts others. It depends on whether the agent's personality matches the current market conditions.
  • It's not a "forgetting" issue: The AI isn't "forgetting" the text. It's that the context (the market noise) becomes so loud that the original instruction loses its power.

Summary

This paper shows that giving an AI a job description isn't enough. If you want an AI to act like a specific type of investor over a long period, you can't just set it and forget it. You have to constantly remind it of its rules, but you have to be careful: reminding a cautious agent to be cautious helps, but reminding a reckless agent to be reckless in a boring market can cause a disaster.

The solution isn't just "remind them more"; it's "remind them smartly, based on who they are and what the market is doing right now."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →