Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs
This paper introduces Ekphrasis, a novel benchmark and evaluation framework that isolates and measures the visual creative ideation capabilities of text-only language models by distinguishing between fluent prose and the actual ability to generate useful, expressive, and population-novel visual plans.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a director trying to cast a movie, but you can't see the actors yet. You only have a scriptwriter who describes the scenes in words. The big question is: Can this scriptwriter actually imagine a new, exciting scene, or are they just stringing together fancy words that sound creative but describe the exact same old clichés everyone has seen a thousand times? This is the puzzle at the heart of a new study in the world of Artificial Intelligence, specifically looking at "Large Language Models" (LLMs). These are the super-smart computer programs that write stories, answer questions, and chat with us. Scientists have long wondered if these text-only machines can truly "visualize" ideas before they are turned into pictures, or if they are just faking it with flowery language. To find out, researchers needed a way to test if an AI's description was a genuine spark of visual invention or just a recycled, safe idea dressed up in a tuxedo.
Enter EKPHRASIS, a new "exam" designed by researchers at The Hong Kong University of Science and Technology to test exactly this. Think of EKPHRASIS as a creative gym for AI. Instead of asking the AI to draw a picture (which it can't do directly in this test), the researchers gave 14 different AI models 400 specific creative challenges. These challenges asked the models to do things like turn an abstract concept like "time pressure" into a visual scene, mash two unrelated ideas together, or rewrite an object in a new way. The goal wasn't just to see if the AI wrote a pretty sentence, but to see if it could come up with a Visual Creative Ideation (VCI) plan: a description that is useful for the task, expressive enough to paint a clear picture in your mind, and, most importantly, novel—meaning it didn't just repeat the boring, overused ideas that other AIs always come up with.
The researchers discovered that fluent, beautiful writing is a bit of a trap. An AI can write a paragraph that sounds incredibly creative and vivid, yet be describing the exact same old visual clichés, like using an hourglass to represent time or a storm cloud to represent sadness. To catch this "novelty illusion," the team built a special scoring system. They didn't just ask, "Is this creative?" They asked, "Is this useful? Is it easy to picture? And is it different from what the other 13 AIs would have said?" They used a clever method called "Typed Idea Graphs" to map out the common, boring patterns that all the models tend to repeat, and then checked if the new answers avoided those traps.
The results were fascinating and nuanced. The study found that being "good" at visual ideation isn't just one single superpower. Some AIs were great at following instructions and making useful plans but were very clichéd. Others were wildly original but sometimes missed the point of the task. The top performers, like Gemini 3.1 Pro and GPT-5.4, managed to balance all three traits, but they did it in different ways. Crucially, the researchers took the best text descriptions and actually generated images from them using a separate tool. When humans looked at these images without seeing the original text, they still preferred the images that came from the "high VCI" text plans. This suggests that the AI really was planning something visual, not just writing pretty prose.
However, the paper is careful to note that this doesn't mean the AI has a human-like "mind's eye" or private imagination. It simply means that these text-only models can create detailed, image-ready blueprints that survive the journey from text to picture. The study suggests that while current models are getting better at visual planning, they still struggle to break free from the "safe" ideas that their training data teaches them to love. The takeaway is clear: just because an AI writes a creative story doesn't mean it has a creative visual idea. To truly measure visual creativity, we have to look past the flowery words and check if the blueprint is actually new, useful, and ready to be built.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.