← Latest papers
📈 economics

Prospect-Theory Behavior from Bellman Optimality in MDPs with Catastrophic States

This paper demonstrates that standard Bellman optimality in Markov decision processes with absorbing catastrophic states inherently generates key prospect-theory signatures—such as S-shaped value functions, endogenous loss aversion, and reflection-effect policy reversals—without requiring any explicit probability weighting, utility curvature, or framing biases.

Original authors: Yujiao Chen

Published 2026-06-02
📖 6 min read🧠 Deep dive

Original authors: Yujiao Chen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are playing a video game where your character has a "Game Over" screen that is permanent. Once you hit that screen, the game ends, and you lose everything. This paper asks a simple question: If you are a perfectly logical, risk-neutral computer trying to win this game, will you naturally start acting like a human who is afraid of losing?

The answer is yes. Even without any built-in fear, anxiety, or "loss aversion" programmed into the brain, the simple math of trying to survive near a "Game Over" boundary creates behavior that looks exactly like the famous psychological patterns known as Prospect Theory.

Here is the breakdown of how this happens, using everyday analogies.

1. The Setup: The Cliff and the Two Paths

Imagine you are standing on a ledge. To your left is a Catastrophe Cliff (falling means instant death/game over). To your right is the rest of the world. You have two choices every turn:

  • The Safe Path: You take a small, guaranteed step forward.
  • The Risky Path: You flip a coin. Heads, you take a giant leap forward. Tails, you take a giant step backward toward the cliff.

Usually, if you are far from the cliff, a logical computer would always choose the Risky Path because, on average, it moves you forward faster. If you are in a "decline" scenario (where both paths move you backward), a logical computer would usually choose the Safe Path to lose slowly.

2. The Surprise: The "S-Shape" and the "Hail Mary"

The paper discovers that when you get close to the cliff, the computer's logic flips completely, creating three strange behaviors that look like human psychology:

A. The "S-Shape" (Playing it Safe when Winning)

When you are doing well (far from the cliff) but getting closer to the edge, the computer suddenly becomes extremely cautious.

  • The Analogy: Imagine you are winning a race and are just 10 meters from the finish line, but there is a giant pit right before the line. Even though taking a huge leap might win you the race faster, you instinctively slow down to a walk. You don't want to trip and fall into the pit.
  • The Result: The computer stops taking risks, even though the risky move has a higher average reward. It prefers a small, safe gain to avoid the tiny chance of falling off the cliff. This creates a value curve that looks like an "S": flat and cautious near the cliff, then curving upward as you get safer.

B. The "Hail Mary" (Gambling when Losing)

Now, imagine you are in a "decline" scenario. Both paths move you backward toward the cliff. The Safe path moves you back slowly; the Risky path moves you back fast, but sometimes jumps you forward.

  • The Analogy: You are losing a football game with 1 second left. The Safe play (kicking a field goal) will lose you the game, but slowly. The Risky play (a long pass) has a high chance of failing, but if it works, you win. A logical computer realizes that "slowly losing" is the same as "losing for sure." So, it decides to gamble everything.
  • The Result: The computer takes the risky action near the cliff, even though it's statistically more likely to lose immediately. It gambles because the only way to escape the cliff is a lucky break.

C. The "Loss Aversion" Illusion

Because of the two behaviors above, the computer starts acting like a human who hates losing twice as much as they love winning.

  • The Illusion: If you measure how much the computer "cares" about a loss versus a gain, the math shows it cares about the loss much more.
  • The Truth: The computer doesn't feel fear. It's just doing the math. The "cost" of falling off the cliff is infinite (game over), so the math forces it to treat a small step backward as a massive threat. This creates a "loss aversion" number that is greater than 1, purely by accident of the environment.

3. The "Magic Formula"

The authors didn't just simulate this; they found a closed-form formula (a simple math equation) that predicts exactly how "loss-averse" the computer will be.

  • The formula depends on three things: how likely you are to win the coin flip, how much you value the future (discount factor), and how much bigger the loss is compared to the win.
  • The Big Discovery: Even if the win and loss are exactly the same size (a perfectly fair coin flip), the mere existence of the "Game Over" cliff makes the computer act like it is loss-averse. The cliff itself is enough to create the behavior.

4. Does it work for real learning agents?

The paper tested this with a "model-free" AI (an agent that learns by trial and error, like a human or a robot, without knowing the rules of the game).

  • The Result: The learning agent figured out the exact same "S-shape" and "Hail Mary" strategies. It didn't need to be told to be safe or desperate; it learned these behaviors naturally just by trying to avoid the cliff.
  • Robustness: This happens even if the game is a bit "noisy" (if the steps aren't perfectly straight, or if there is random shaking). The behavior holds up.

Summary

This paper argues that Prospect Theory (the idea that humans are irrational and fear losses more than they love gains) might not always be a psychological flaw. Instead, it might be a rational survival strategy for any intelligent agent facing a "Game Over" scenario.

If you are near a catastrophic failure, the smartest thing to do is:

  1. Stop gambling when you are winning (to protect your lead).
  2. Start gambling when you are losing (because you have nothing to lose).

The paper shows that a perfectly logical, unfeeling machine will naturally adopt these exact strategies, making it look like it has human psychology, even though it's just doing the math.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →