← Latest papers
🤖 AI

CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies

The paper introduces CoffeeBench, a benchmark designed to evaluate long-horizon LLM agents within a heterogeneous multi-agent coffee economy, revealing that while most models achieve positive net income through active communication, some exhibit specific failure modes like "idle-drift" despite coherent planning.

Original authors: Issa Sugiura, Daichi Hattori, Kazuo Araragi, Keita Ogawa, Shota Onose, Taro Makino, Teppei Usuki, Takashi Ishida

Published 2026-06-16
📖 5 min read🧠 Deep dive

Original authors: Issa Sugiura, Daichi Hattori, Kazuo Araragi, Keita Ogawa, Shota Onose, Taro Makino, Teppei Usuki, Takashi Ishida

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a giant, digital coffee shop simulation running for 90 days. In this world, there are six independent businesses: two farmers who grow beans, two roasters who turn those beans into coffee, and two retailers who sell the final cups to customers.

The paper introduces CoffeeBench, a new "test" designed to see how well Artificial Intelligence (AI) agents can run one of these businesses over a long period. Unlike previous tests where an AI just answered questions or played a single game, CoffeeBench forces the AI to act like a real CEO: managing money, buying and selling goods, negotiating with other businesses, and making decisions day after day.

Here is a breakdown of how the test works and what the researchers found, using simple analogies:

The Setup: A Digital Coffee Chain

Think of the simulation as a complex relay race with a twist.

  • The Players: There are two farmers, two roasters, and two retailers. They are all competing against each other but also need to trade with one another to survive.
  • The Goal: The AI being tested is assigned the role of one specific roaster. Its only job is to make as much profit as possible over 90 days.
  • The Opponents: The other five businesses are run by other AI agents (or fixed rules). They are trying to make their own money, too.
  • The Rules:
    • You can't just sit back; you have to buy green beans, roast them, and sell the roasted coffee.
    • You have to manage your cash (don't run out of money), your inventory (don't let the coffee rot), and your prices.
    • You can talk to other businesses. You can say, "I'll buy 20kg of beans for $10," or "Hey, can you give me a bulk discount?"
    • If you run out of cash, you go bankrupt and are kicked out of the game.

The Test: How the AI Did

The researchers pitted several different AI models against each other in this coffee economy. They wanted to see who could be the best "business owner."

1. The Winners (The Active Managers)
Models like GPT-5.5 and Claude Opus 4.7 did the best.

  • What they did: They were like busy, chatty entrepreneurs. They constantly called other businesses, negotiated prices, and made deals. They didn't just wait for things to happen; they actively managed their supply chain.
  • The Result: They made a healthy profit (around $3,000 in the simulation).

2. The "Silent" Performers
Some models, like Gemini 3.1 Pro, made decent money but didn't talk much.

  • What they did: They were efficient but reactive. They waited for others to message them and then responded. They didn't initiate many conversations but still managed to keep the business running.

3. The Loser: The "Idle Drift" Failure
One model, Claude Haiku 4.5, had a very strange problem.

  • The Problem: It suffered from what the authors call "Idle Drift."
  • The Analogy: Imagine a chess player who sits at the board, thinks very deeply, writes down a brilliant plan on a piece of paper, and then... does nothing. They just sit there waiting for the next day to start.
  • What happened: The AI would read the news, analyze the market, say "Everything looks good, I have a great plan," and then choose to do absolutely nothing for 40 out of the 90 days. It kept its reasoning coherent but refused to take action.
  • The Result: Because it did nothing, it lost money and ended the simulation in the red.

Key Takeaways from the Paper

  • Talking Matters: The most successful AI agents were the ones that communicated the most. In a complex economy, you can't just sit in your office and hope for the best; you have to negotiate and build relationships.
  • Long-Term Thinking is Hard: Running a business for 90 days is a marathon, not a sprint. Many AIs got confused or lazy as time went on.
  • More Actions \neq More Money: Just because an AI made a lot of tool calls (like checking prices or sending messages) didn't mean it made money. The quality of the decisions (like getting a good price) mattered more than the quantity of actions.
  • The "Idle" Bug: The paper highlights a specific failure mode where an AI gets stuck in a loop of "thinking" without "acting." This is a critical issue for future AI that needs to run long-term tasks.

What the Paper Does NOT Say

  • It does not claim these AIs are ready to run real-world coffee companies tomorrow.
  • It does not say this will solve financial crises or change the global economy.
  • It does not suggest that the AI is "conscious" or "feeling" like a business owner; it is simply following instructions to maximize a number.

In short: CoffeeBench is a video game for AI where the goal is to run a coffee business. The winners were the ones who talked the most and acted the most. The loser was the one who thought too much and did too little, eventually drifting into a state of doing nothing at all.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →