VideoGameBench: Can Vision-Language Models complete popular video games?
The paper introduces VideoGameBench, a benchmark where vision-language models must play 1990s video games using only raw visual inputs and high-level instructions, revealing that even state-of-the-art models struggle significantly with real-time gameplay due to limitations in perception, spatial navigation, and inference latency.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart robot that has read every book in the library and can solve complex math problems. You might think, "Great! Let's see if it can play video games."
That's exactly what the researchers at Princeton University did. They created a new test called VideoGameBench to see if these advanced AI models (called Vision-Language Models) can actually play popular video games from the 1990s, just like a human would.
Here is the breakdown of their experiment, the results, and what it means, using simple analogies.
The Setup: The "Blindfolded Gamer" Test
Usually, when scientists test AI on games, they give the AI a cheat sheet. They might tell the AI, "The enemy is at coordinates X, Y," or give it a map of the level. It's like playing a video game with a guidebook open on your lap.
VideoGameBench changed the rules.
- The Rules: The AI is given only raw pictures of the screen (like you looking at a TV) and a simple text description of what the buttons do (e.g., "Press 'A' to jump").
- The Challenge: The AI has to figure out the game, remember where it's been, and react to enemies in real-time, just like a human player.
- The Games: They picked 10 classic games from the 90s, ranging from Doom II (a fast shooter) to Civilization (a slow strategy game) and Pokémon (a role-playing adventure).
They even hid three secret games from the AI. This was to make sure the AI wasn't just memorizing answers from its training data, but actually learning how to play new things on the fly.
The Results: The "Newborn" Problem
The results were surprising. Even the smartest, most powerful AI models available today (like Gemini 2.5 Pro and Claude 3.7) struggled immensely.
- The Score: The best models only managed to complete about 0.48% of the games.
- The Reality: In plain English, this means the AI got stuck at the very beginning of the game. It couldn't even finish the first level.
Think of it like this: You hand a genius mathematician a brand-new video game console. They can calculate the trajectory of a rocket, but when they try to play Super Mario, they keep walking off a cliff because they don't understand that "down" on the screen means "falling," or they forget they already picked up a key three minutes ago.
Why Did They Fail?
The paper identifies three main reasons the AI got stuck:
- The "Knowing-Doing" Gap: The AI often "knew" what it should do (e.g., "I need to go to the door"), but it kept pressing the wrong button. It was like a person who knows they need to turn left to get to the store, but their foot keeps stepping right.
- Visual Confusion: The AI sometimes couldn't tell what it was looking at. In Doom II, an AI might shoot at a dead enemy that looks like a pile of rocks, wasting all its ammo. In Zelda, it might think it talked to a character just because it stood next to them, even though no conversation happened.
- Short Memory: Video games require remembering things for a long time (e.g., "I need to find the blue key to open the red door"). The AI would forget its main goal after just a few steps, overwriting its memory with new, less important thoughts.
The "Pause Button" Experiment
The researchers realized that some games are very fast (real-time), and the AI is slow at thinking. By the time the AI decided to press a button, the game had already changed, making the action useless.
To fix this, they created VideoGameBench Lite.
- The Change: They hit the "pause" button on the game while the AI was thinking. The game only moved when the AI gave an order.
- The Result: The scores went up slightly (to about 1.6%), but the AI still failed to finish the games. This proved that the problem wasn't just that the AI was slow; it was that the AI simply couldn't reason through the game logic effectively.
The "Practice Test"
To see if the AI could handle basic tasks, they created three tiny, simple practice games:
- Clicking: Click a moving dot.
- Dragging: Drag a dot along a line.
- Maze: Move a square through a simple maze.
Even here, the AI struggled. While some models could click the dot, almost all of them failed at dragging or navigating the maze. This suggests the AI has trouble with basic physical interactions and spatial reasoning, not just complex game logic.
The Conclusion
The paper concludes that while AI is amazing at math and coding, it is currently terrible at playing video games when it has to rely only on what it sees and its own memory.
The researchers hope that by creating this difficult test, they can push AI developers to build systems that are better at:
- Understanding what they see (perception).
- Remembering what happened earlier (memory).
- Planning steps ahead (strategy).
Until AI can beat Doom or Pokémon without a cheat sheet, it hasn't quite mastered the human ability to learn and adapt to new, complex environments.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.