← Latest papers
📈 economics

Bridging Behavioral Finance and Robust Agentic DRL: A novel computational framework for Portfolio Selection

This study introduces RPPO-MPT, a novel robust reinforcement learning framework that integrates perturbation-based noise reduction with Prospect Theory-inspired reward shaping to outperform conventional models in equity portfolio optimization by better aligning with investor behavioral preferences under market uncertainty.

Original authors: Ehsan Ameri, Majid Mirzaee Ghazani, Donya Rahmani

Published 2026-06-26
📖 4 min read☕ Coffee break read

Original authors: Ehsan Ameri, Majid Mirzaee Ghazani, Donya Rahmani

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are the captain of a ship navigating a very stormy ocean. The ocean represents the stock market, and your ship is your investment portfolio.

For a long time, captains (investors) have used two main ways to steer:

  1. The Old Map (Traditional Finance): This assumes the ocean is predictable and that sailors are perfectly logical robots who only care about reaching the destination as fast as possible, regardless of how scary the waves get.
  2. The New GPS (Standard AI): This uses computers to learn from past storms. However, most of these AI captains still act like robots. They get excited by big waves (gains) but panic just as easily as humans do when the water gets choppy, often making decisions that don't feel "right" to a real human investor.

This paper introduces a new, super-smart AI captain called RPPO-MPT. It's designed to handle the stormy ocean better by combining three specific tricks.

The Three Tricks of the New Captain

1. The "What If?" Training (Robustness)
Imagine you are practicing for a race. A normal captain practices on a calm day. If a sudden storm hits on race day, they might freeze.
The new captain, however, practices in a simulator where the computer intentionally adds fake storms, wind gusts, and fog to the training data. It asks, "What if the wind blows 5% harder? What if the current shifts slightly?"
By training in these "fake" chaotic conditions, the captain learns to stay calm and steer correctly even when the real ocean gets messy. This is called perturbation-based robust learning. It makes the AI less sensitive to noise and less likely to make a panic decision when the market jumps around.

2. The "Human Heart" Reward System (Prospect Theory)
Standard AI captains get a "gold star" (reward) for every dollar they make. But real humans don't feel that way.

  • The Fear: Losing $100 feels much worse than the joy of finding $100 feels good.
  • The Hope: We get excited about gains, but we are terrified of losses.
    The new captain is programmed with a "Hope-Fear" reward system. Instead of just counting dollars, it calculates a score based on how a human would feel.
  • If the ship makes money, it gets a "Hope" score (but the excitement grows slower as you get richer).
  • If the ship loses money, it gets a "Fear" score (and the pain is multiplied by 2.25 times!).
    This teaches the AI to be naturally more careful about losing money than it is eager to make it, mimicking how real people actually behave.

3. The Hybrid Approach
The paper compares three versions of the captain:

  • The Basic Captain (PPO): Uses the standard AI GPS. Good, but can be shaken by storms.
  • The Tough Captain (RPPO): Uses the "What If?" training. Very steady, but doesn't understand human fear.
  • The Super Captain (RPPO-MPT): Uses both the "What If?" training AND the "Human Heart" reward system.

The Race Results

The authors tested these captains using data from 15 big American companies (like Apple and Amazon) over a long period (2010 to 2025). They split the time into two parts: a "practice" period (2010–2022) and a "real race" period (2022–2025).

Here is what happened:

  • The Basic Captain did okay, but sometimes got shaken up by the market's wild swings.
  • The Tough Captain was steadier and handled the noise well.
  • The Super Captain (RPPO-MPT) won the race.

Why did the Super Captain win?

  • Higher Returns: It made more money over time.
  • Better Safety: It didn't lose as much money during the scary parts of the race (lower "drawdown").
  • Happier Investor: Because it understood the "Hope-Fear" balance, it produced results that felt better to a human investor, offering a better "utility" score.

The Bottom Line

This paper claims that by teaching an AI to practice in fake chaos (to build toughness) and think like a human (to understand fear and hope), you can create a portfolio manager that is both safer and more profitable than standard AI models. It bridges the gap between cold, hard math and the messy, emotional reality of how people actually invest.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →