StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley
StarDojo is a novel benchmark based on the game Stardew Valley that evaluates open-ended agentic multimodal LLMs in complex production-living simulations through 1,000 curated tasks, revealing that even state-of-the-art models like GPT-4.1 struggle significantly with visual understanding, reasoning, and low-level manipulation, achieving only a 12.7% success rate.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to live a full human life. Not just how to answer a question or write a poem, but how to wake up, go to work, make friends, manage a budget, and handle unexpected rainstorms—all at the same time.
That is exactly what the authors of this paper, StarDojo, are trying to do. They have built a testing ground using the popular video game Stardew Valley to see if Artificial Intelligence (AI) agents can handle the messy, complicated reality of "production and living."
Here is a breakdown of their work using simple analogies:
1. The Problem: The "Two-Track" Gap
Think of current AI benchmarks like a driving test where you only have to park a car in a straight line (production) or a test where you only have to chat with a passenger (social interaction).
- The Reality: In real life, you can't just park and chat. You have to drive to the grocery store (production), buy food, and then talk to the cashier (social) while managing your time and money.
- The Gap: Existing tests don't check if an AI can do both at once. They are usually too simple or too focused on just one thing.
2. The Solution: StarDojo (The "Life Simulator" Exam)
The authors created StarDojo, a new benchmark based on the game Stardew Valley.
- The Game: Imagine a cozy farm where you plant crops, mine for rocks, fight slimes in a cave, and hang out with villagers.
- The Twist: They turned this game into a rigorous exam for AI. Instead of a human playing, they let AI agents try to complete specific tasks.
- The Tasks: They created 1,000 different challenges ranging from "easy" (like watering a plant) to "hard" (like planning a whole season of farming while building a relationship with a neighbor and fighting monsters).
- The "Lite" Version: To make it easier to test, they also picked a smaller set of 100 representative tasks (StarDojo-Lite) that cover the basics.
3. How They Connected the AI to the Game
Normally, to control a video game, you need a keyboard and mouse. But you can't easily plug a keyboard into a robot brain.
- The Bridge: The authors built a special "translator" (called StarDojoMod) that lets the AI talk directly to the game's engine.
- The Analogy: Instead of the AI trying to "see" the screen and "press" buttons like a human, it's like the AI has a direct phone line to the game's internal computer. It can ask, "What is my energy level?" or say, "Move me to the left," and the game does it instantly. This allows them to run many games at the same time to test the AI quickly.
4. The Results: The AI Struggles to "Grow Up"
The authors tested the smartest AI models available (like GPT-4.1, Claude, and Gemini) on this exam. The results were sobering:
- The Score: Even the best AI only got about 12.7% of the tasks right.
- The Easy Stuff: The AI was okay at simple, short tasks (like picking up a tool if it was already in its hand).
- The Hard Stuff: The AI completely failed at anything requiring a long plan or moving around the map. If a task took more than a few steps, the AI got lost or forgot what it was doing.
5. Why Did the AI Fail? (The "Four Blind Spots")
The authors analyzed the mistakes and found four main reasons the AI couldn't handle this "life simulator":
- Visual Blindness (42% of errors): The AI looked at the game screen and couldn't tell what was what. It might think a rock is a tree, or it can't see that a door is right in front of it. It's like trying to navigate a room while wearing foggy glasses.
- Reasoning Confusion (21%): Even when the AI had the right information (like a text list saying "You are near the barn"), it ignored the text and relied on its blurry vision, leading to confusion.
- Short Attention Span (21%): The AI couldn't plan ahead. It might forget that it needs to water the crops before it can harvest them, or it would switch goals in the middle of a task.
- Clumsy Hands (16%): The AI knew what to do but messed up the how. It might try to use a tool in the wrong direction or pick the wrong item from its inventory.
6. The Takeaway
The paper concludes that while AI is great at answering questions or writing code, it is currently very bad at embodied living. It struggles to combine seeing, thinking, planning, and acting in a dynamic world where time passes and things change.
StarDojo is now an open tool for researchers to use. It's like a "gym" where AI can train to get better at the complex, messy, and beautiful challenge of living a simulated life. The paper doesn't promise that AI will be running farms tomorrow; it just shows us exactly where the AI is stumbling so we can help it learn.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.