← Latest papers
🤖 AI

GEBench: Benchmarking Image Generation Models as GUI Environments

GEBench is a new benchmark and a multi-dimensional metric (GE-Score) designed to evaluate the ability of image generation models to produce temporally coherent and logically consistent GUI state transitions across single-step and multi-step interactions.

Original authors: Haodong Li, Jingwei Wu, Quan Sun, Guopeng Li, Juanxi Tian, Huanyu Zhang, Yanlin Lai, Ruichuan An, Hongbo Peng, Yuhong Dai, Chenxi Li, Chunmei Qing, Jia Wang, Ziyang Meng, Zheng Ge, Xiangyu Zhang, Daxi
Published 2026-02-11
📖 4 min read☕ Coffee break read

Original authors: Haodong Li, Jingwei Wu, Quan Sun, Guopeng Li, Juanxi Tian, Huanyu Zhang, Yanlin Lai, Ruichuan An, Hongbo Peng, Yuhong Dai, Chenxi Li, Chunmei Qing, Jia Wang, Ziyang Meng, Zheng Ge, Xiangyu Zhang, Daxin Jiang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are playing a video game, but instead of a pre-made world, the game is being "dreamed up" in real-time by an AI. You press a button to "Open Settings," and the AI instantly paints a brand-new picture of what that settings menu should look like.

This sounds magical, but there is a huge problem: How do we know if the AI is actually "playing the game" correctly, or if it’s just hallucinating random, beautiful nonsense?

That is exactly what this paper, GEBench, is about. Here is the breakdown:

1. The Problem: The "Beautiful Liar" Paradox

Current AI models (like those that make art) are amazing at making pretty pictures. However, if you ask an AI to show you a "Phone Settings" screen and then ask it to "Click the Wi-Fi button," the AI might show you a beautiful picture of a phone, but the Wi-Fi button might be in the wrong place, the text might look like gibberish, or the next screen might look like a completely different phone.

In the AI world, we call this a lack of temporal coherence and spatial grounding. In everyday language, it means the AI has a "short-term memory problem" and "bad hand-eye coordination." It can draw a single frame perfectly, but it can't follow a logical story.

2. The Solution: GEBench (The Ultimate "UI Driving Test")

The researchers created GEBench, which acts like a rigorous driving test for AI. Instead of just asking the AI to "draw a cat," they give it specific "driving" tasks within a digital interface:

  • The Single Step: "Click this specific icon. Show me what happens next." (Testing basic reflexes).
  • The Long Journey: "Start at the home screen and successfully order a coffee." (Testing long-term planning and memory).
  • The Imaginary App: "Design a brand new app for a space traveler." (Testing creativity and logic).
  • The Precision Test: "Click exactly at these coordinates [X, Y]." (Testing hand-eye coordination).

3. The Scoring: The "GE-Score" (The Five-Star Review)

To grade the AI, they don't just say "looks good" or "looks bad." They use a five-dimensional grading system, much like a restaurant critic:

  1. Goal Achievement (Did you get the food?): Did the AI actually do what you asked? If you asked to open a menu, is the menu open?
  2. Interaction Logic (Does the kitchen make sense?): If you click "Delete," does the app actually show a "Are you sure?" pop-up, or does it just teleport to a random screen?
  3. Consistency (Is the waiter the same person?): If you move from one screen to the next, does the phone still look like the same phone, or did the color and shape suddenly change?
  4. UI Plausibility (Is the menu readable?): Do the buttons look like real buttons, or do they look like weird, melted blobs?
  5. Visual Quality (Is the presentation pretty?): Is the text sharp and clear, or is it blurry and messy?

4. The Verdict: The AI is a "Talented but Distracted Student"

After testing the world's best AIs (like Google's and OpenAI's models), the researchers found something interesting:

The AIs are brilliant at single steps (they can draw a single, beautiful screen), but they fail miserably at long sequences. As the "story" goes on, the AI gets confused. It forgets where it was, it loses track of the coordinates, and it starts "hallucinating" text and icons.

Why does this matter?

In the future, we want AI "Agents" that can actually use your computer for you—booking flights, organizing files, or managing your emails. To do that, the AI needs to live in a reliable "world." GEBench is the yardstick that will help scientists build an AI that doesn't just dream of interfaces, but actually understands how to navigate them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →