← Latest papers
🤖 AI

Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions

This paper introduces Mini Amusement Parks (MAPs), a novel amusement-park simulator designed to holistically benchmark AI agents on complex, real-world business decision-making tasks, revealing that current state-of-the-art LLMs significantly underperform human baselines in long-horizon optimization, sample-efficient learning, spatial reasoning, and world modeling.

Original authors: Stéphane Aroca-Ouellette, Ian Berlot-Attwell, Panagiotis Lymperopoulos, Abhiramon Rajasekharan, Tongqi Zhu, Herin Kang, Kaheer Suleman, Sam Pasupalak

Published 2026-07-14
📖 6 min read🧠 Deep dive

Original authors: Stéphane Aroca-Ouellette, Ian Berlot-Attwell, Panagiotis Lymperopoulos, Abhiramon Rajasekharan, Tongqi Zhu, Herin Kang, Kaheer Suleman, Sam Pasupalak

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you've just been handed the keys to the world's most chaotic, unpredictable, and profitable amusement park. Your job? Be the manager. You have to decide where to put the roller coasters, how many ice cream stands to open, when to hire a janitor, and whether to spend your cash on researching a new, super-expensive ride. But here's the catch: the guests are moody, the rides break down randomly, and if you make the wrong move, your park could go bankrupt before lunch.

This is the world of Mini Amusement Parks (MAPS), a new video game-like simulator created by researchers to test how well Artificial Intelligence (AI) can run a real business.

The Big Test: AI vs. Human Managers

The researchers set up a showdown between the smartest AI models on the planet (like GPT-5, Claude, and Gemini) and expert human players. The goal was simple: see who could build the most valuable park.

The results were a massive wake-up call. Even the best AI models struggled to keep up.

  • On the Easy mode, the top AI (GPT-5) managed to reach only 8.80% of the score that expert humans achieved. That means the humans were 11.4 times better.
  • On the Medium mode, where the game gets trickier and requires long-term thinking, the gap widened. The AI's performance dropped to just 4.44% of the human score, making the humans 15.3 times superior.

It's like giving a Formula 1 car to a toddler and a professional racer to a pro; the toddler (the AI) barely gets the car moving, while the pro (the human) is lapping the track.

Why Did the AI Fail?

The paper suggests the AI isn't just "bad at math"; it's missing some very human superpowers. The researchers identified five specific areas where the AI stumbled:

  1. The "What's for Lunch?" Problem (Short-Sightedness):
    Humans plan ahead. They know that if they build a giant roller coaster today, they need to hire more staff and save money for next month. The AI, however, tends to be myopic. It grabs the immediate reward (like building a ride) without thinking about the consequences. It's like a kid eating all the candy now and getting a stomach ache, forgetting they need energy for the rest of the day. When the game got harder (Medium mode), the AI's performance tanked because it couldn't see past the next turn.

  2. The "Blindfolded" Problem (Spatial Reasoning):
    In a real park, where you put a shop matters. You want it near a path, not stuck in a corner. The AI often placed rides in spots that were hard to reach or clustered them all together, ignoring the flow of the crowd. Interestingly, giving the AI a picture of the park (vision) didn't help; in fact, it sometimes made things worse. The AI seemed confused by the visual data, whereas a simple set of rules (a "heuristic") for where to place things actually helped the AI perform better.

  3. The "Guessing Game" Problem (Stochasticity):
    Real life is messy. Sometimes a ride breaks down for no reason; sometimes a guest is just grumpy. This is called "stochasticity." The AI struggled to handle this randomness. It couldn't tell the difference between a bad day caused by bad luck and a bad day caused by a bad decision. When the researchers tried to use the AI to predict the future (a "world model"), it often made things worse, leading to slower and more expensive decisions.

  4. The "Amnesia" Problem (Active Learning):
    Humans are great at experimenting. If you try a new strategy and it fails, you learn why and try something else. The researchers gave the AI a "sandbox" mode—a safe practice area where it could experiment for 100 days before the real test. Surprisingly, this didn't help much. The AI mostly just repeated what was already in the instruction manual or made notes that were too specific to be useful later. It failed to build a true "mental model" of how the park worked.

  5. The "Research Rabbit Hole" Problem:
    In the harder version of the game, you have to research new rides before you can build them. The AI often picked the most expensive, flashy rides to research immediately, even when it couldn't afford to build them yet. It was like a student trying to buy a Ferrari before they even have a driver's license.

What the Paper Rules Out

The paper is very clear about what doesn't work:

  • Just giving the AI a picture isn't the fix. Adding visual input to the AI's "brain" didn't solve the spatial problems; in fact, it often confused the models more than just text alone.
  • Current "World Models" aren't ready. The researchers tried using a specific AI technique (WALL-E) to help the AI predict the future, but it actually made the AI perform worse and cost much more to run.
  • Practice doesn't always make perfect. Simply letting the AI play in a sandbox for 100 days didn't automatically make it smarter. In many cases, the AI just got worse or stayed the same because it couldn't figure out what to learn from its mistakes.

How Sure Are We?

These findings are based on simulations and measurements within the MAPS game environment. The researchers ran the AI and human players through specific layouts (maps) multiple times to get these numbers. They didn't just guess; they measured the park value, the money spent, and the time taken.

The paper suggests that while AI is getting better at many things, it still has a long way to go before it can handle the messy, interconnected, and unpredictable nature of running a real business. The gap between human experts and AI isn't just a little bit; it's a chasm. The researchers hope that by using this "Mini Amusement Park" as a testbed, we can finally start building AI that can think, plan, and adapt like a true human manager.

Until then, if you want to run a theme park, you might want to stick with a human manager—and maybe keep the AI on a very short leash.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →