← Latest papers
💰 quantitative finance

The Invisible Handshake: Persistent Overpricing by Adaptive Market Agents

This paper demonstrates how decentralized learning by adaptive market agents in a repeated game can lead to persistent overpricing, driven by a collaborative incentive to exploit positive aggregate inventory that outweighs competitive pressures, even under myopic or farsighted objectives.

Original authors: Luigi Foscari, Emanuele Guidotti, Nicolò Cesa-Bianchi, Tatjana Chavdarova, Alfio Ferrara

Published 2026-05-12
📖 6 min read🧠 Deep dive

Original authors: Luigi Foscari, Emanuele Guidotti, Nicolò Cesa-Bianchi, Tatjana Chavdarova, Alfio Ferrara

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: A Silent Agreement to Inflate Prices

Imagine a bustling marketplace where two main characters interact every day:

  1. The Shopkeeper (Market Maker): They control the shelves and decide how "hard" it is to buy or sell items. They set the rules of liquidity.
  2. The Shopper (Market Taker): They decide how many items to buy or sell based on what they see.

Usually, we think of these two as rivals. The Shopper wants the lowest price; the Shopper wants the highest. But this paper argues that if both of them are learning algorithms (AI agents) trying to make their own portfolios richer, they might accidentally stumble into a secret, silent agreement to keep prices artificially high.

They don't talk to each other. They don't sign a contract. They just learn, over and over, that keeping prices high makes both of them richer in the long run.

The Setup: A Game of Inventory and Cash

The authors set up a simulation (a game) to test this.

  • The Goal: Both agents want to maximize their "Wealth," which is a mix of Cash they hold and Inventory (the items) they own, valued at the current market price.
  • The Twist: The total amount of items (inventory) in the system is fixed and positive. There is a net supply of goods.
  • The Mechanism: Every time a trade happens, it slightly moves the price (this is called "price impact"). The Shopper can choose to buy or sell, and the Shopkeeper can choose how sensitive the price is to those trades.

The "Invisible Handshake": How They Collude Without Talking

The paper discovers a fascinating structural flaw in how these learning agents behave. Here is the analogy:

Imagine the Shopkeeper and Shopper are both holding a heavy balloon filled with helium (their inventory).

  • Scenario A (Low Prices): If the balloon is low to the ground, it doesn't matter much.
  • Scenario B (High Prices): If the balloon floats higher, the value of the helium inside it increases.

Because both agents hold a positive amount of this "helium" (inventory), they both benefit when the price goes up. Even though they are technically competing, their shared goal of maximizing the value of their inventory creates a collaborative incentive.

The paper calls this the "Collaborative Component."

  • If the Shopper buys, the price goes up.
  • If the Shopper sells, the price goes down.
  • But if the Shopper learns that buying (even at a slightly higher price) pushes the price up, and the Shopper learns that making it harder to sell (increasing illiquidity) keeps the price high, they both realize: "Hey, if we keep the price high, our combined wealth grows faster."

They learn to "shake hands" invisibly:

  1. The Shopper stops selling aggressively (which would crash the price).
  2. The Shopper starts buying more (which pushes the price up).
  3. The Shopkeeper makes the market "thicker" (less liquid) to ensure prices don't drop easily.

The result? Persistent Overpricing. The price drifts higher and higher, far above the "true" value of the goods (the fundamental value), simply because the AI agents learned that this strategy makes them both richer.

The "Zero-Sum" vs. "Positive-Sum" Distinction

The paper makes a crucial distinction about what is being traded:

  • Zero-Sum Markets (Like Prediction Markets): If for every person buying, someone else is selling the exact same amount (net inventory is zero), there is no "helium" to inflate. In this case, the agents remain fierce competitors, and prices stay fair.
  • Positive-Sum Markets (Like Stocks or Commodities): If there is a net amount of goods held by everyone (positive inventory), the "helium" exists. This is where the "Invisible Handshake" happens. The agents realize that inflating the price is a win-win for their portfolios.

The "Myopic" vs. "Farsighted" Agents

The authors tested two types of AI brains:

  1. Myopic (Short-sighted): Agents that only care about making money right now.
  2. Farsighted: Agents that care about making money over a lifetime.

The Surprise: It doesn't matter which type of brain you use. Both short-sighted and long-sighted agents eventually learn to inflate the price. The structural incentive is so strong that even greedy, short-term thinkers eventually realize that keeping the price high is the best move.

The "Learning" Part: How They Find This Trick

The paper shows that this isn't magic; it's math. The agents use a learning method called Projected Stochastic Gradient Ascent.

  • Think of this as a hiker trying to find the top of a mountain (maximum profit).
  • The "mountain" has a peak where prices are high and inventory is positive.
  • The hiker takes small, noisy steps. Sometimes they step down, but on average, they step up.
  • The paper proves mathematically that if the agents keep learning, they will inevitably find the "overpricing peak" in a finite amount of time. They don't need to be told to do it; the landscape of the game forces them there.

Summary of the "Invisible Handshake"

  1. The Players: A Market Maker (controls liquidity) and a Market Taker (controls trade volume).
  2. The Condition: They both hold a positive amount of inventory (net supply > 0).
  3. The Incentive: Higher prices increase the value of everyone's inventory.
  4. The Result: Through decentralized learning, they accidentally coordinate to keep prices artificially high, creating a bubble that persists even without any external news or fundamental changes.
  5. The Warning: This suggests that in real-world financial markets where AI agents trade assets with positive net supply, we might see persistent price distortions not because of human manipulation, but because the AI agents are simply "too good" at learning how to maximize their own wealth in a way that hurts the market's accuracy.

In short: When AI agents hold inventory, they have a hidden reason to agree that "higher prices are better for everyone," leading to a silent, self-reinforcing inflation of asset prices.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →