Benchmarking Frontier Text-to-Image Models on Image-Description Prompts
This paper benchmarks four frontier text-to-image models (Gemini 3 Pro Image, FLUX.2, Ideogram 3.0, and Hunyuan 3.0) on 48 highly complex prompts, revealing that Gemini 3 Pro Image leads with a score of 84.8/100 while highlighting that current systems primarily struggle with object counting, geometric accuracy, and text legibility.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the quiet hum of modern computing, a new kind of artist has emerged, one that paints not with brushes or clay, but with words. These are text-to-image systems, computer programs that listen to a human description and attempt to draw a picture that matches it. For years, these systems have been tested on simple requests, like "a cat sitting on a mat," where they perform with surprising grace. But the real world is rarely so simple. When a user asks for something specific and complex—such as a scene with exactly four red chairs, a sign that reads "OPEN" in clear letters, and a reflection in a puddle that obeys the laws of physics—the programs often stumble. They might draw three chairs instead of four, scramble the letters on the sign, or make the reflection float in the air. Understanding why these systems fail, and which ones fail the least, is crucial for anyone hoping to use them for serious work, from designing advertisements to creating storyboards. The gap between a system that works well on easy tasks and one that can handle the messy, detailed demands of reality is where the true test of intelligence lies.
A team of researchers set out to measure this gap by putting four of the most advanced image-generating systems to a rigorous test. They did not ask the computers to draw simple pictures; instead, they fed them the forty-eight most difficult descriptions available in a large collection of real-world image captions. These prompts were chosen because they demanded precision: exact counts of objects, specific spatial arrangements, and legible text embedded within the scene. The systems tested were Hunyuan 3.0, Gemini 3 Pro Image, Black Forest Labs FLUX.2, and Ideogram 3.0. To ensure a fair and unbiased judgment, the researchers employed a clever two-step process. First, one artificial intelligence model acted as a writer, creating a detailed checklist of what a perfect image should look like for each specific prompt, while also looking at the actual image the system had produced to note any obvious flaws. Then, a completely different artificial intelligence model, acting as a judge, used that checklist to grade the image. This separation ensured that the model deciding what to look for was never the same one deciding if the image passed the test, removing the bias that can occur when a system grades its own work.
The results revealed a clear divide in capability. Two systems, Gemini 3 Pro Image and FLUX.2, emerged as the leaders, scoring 84.8 and 82.3 out of 100, respectively. They were followed by a significant gap, with Ideogram 3.0 scoring 65.7 and Hunyuan 3.0 scoring 63.3. The difference between the top and bottom performers was not just a matter of small errors; it was a fundamental difference in how they handled complexity. The leading systems generally got the big picture right but occasionally miscounted objects or introduced minor geometric glitches, such as a wheel merging strangely with a frame. In contrast, the trailing systems struggled with the basics. They frequently failed to include requested elements entirely, such as omitting a specific object from the scene, or they garbled the text, rendering words as unreadable scribbles instead of clear letters. For instance, when asked to draw a mailbox shaped like a car with a specific address on it, the top system got every detail correct, while the lowest-scoring system failed to make the mailbox look like a car at all, missed the address, and drew only two tires instead of four.
The study highlights that while these systems are improving, they still lack a native understanding of discrete, symbolic rules. They are excellent at capturing the general mood, color, and composition of a scene, but they struggle with the precise logic required to count items or spell words correctly. The researchers found that the hardest challenges for every system were exact sequential labeling, such as numbering starting blocks from five to eight, and counting specific numbers of people or objects. Even the best-performing system, Gemini 3 Pro Image, lost points on these tasks, though it did so far less often than the others. The trailing systems, particularly Ideogram 3.0, showed a pattern of omitting entire requested elements, suggesting they sometimes give up on complex instructions rather than attempting to get the details wrong. This distinction is vital for users: if a project requires a scene with many distinct objects, the leading systems are far more reliable, but if the goal is to generate text within an image, even the best systems still struggle to get it right every time.
Ultimately, this research demonstrates that the ability to follow complex instructions is the new frontier for image generation. The gap between the best and the rest is not a matter of style or artistic flair, but of logical consistency. The two leading systems proved they can handle the heavy lifting of compositional demands, while the others remain prone to fundamental errors that break the illusion of reality. As these tools move from experimental curiosities to production tools, the choice of which system to use will depend heavily on the specific task. For simple, atmospheric images, the differences may be negligible, but for requests that demand precision, the leaders offer a level of reliability that the others have yet to match. The path forward involves refining these systems to better understand the rules of counting, text, and space, turning them from artists who guess at details into tools that can execute them with certainty.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.